Control plane for agents & engineers to provision compute and run training & inference across NVIDIA, AMD, and other chips — on clouds, Kubernetes, and on-prem.dstack.ai github.com/dstackai/dstackJoined January 2020
Going to @PyTorch Conference in San Jose next month? @CrusoeAI and we are hosting an evening on day one, October 20, 6:30 PM, right after the sessions.
We're gathering GPU experts, AI researchers, AI infra engineers, and anyone into AI infra and heterogeneous AI compute.
Space is limited: luma.com/6j69xyal
dstack 0.21.3 is out!
Presets: optimized inference configurations can now be pushed to a private registry and later deployed to any cloud, Kubernetes, or on-prem cluster.
Backends: @seeweblive GPU cloud is now supported.
Plus improvements to services and gateways. Release notes: github.com/dstackai/dstac…
Here's one example of using the presets toolkit to optimize Qwen3.8-27B on a single @AMD MI300X.
Through compound learning across linked sessions and source-level patches, performance improved from 311 to 495 tok/s: +59%.
The resulting preset can be deployed on any AMD cloud, Kubernetes cluster, or bare-metal fleet.
Inference serving is increasingly open source. But inference optimization still happens inside each provider, and the optimized deployment remains tied to its proprietary stack. This has to change.
Introducing Presets: an open-source toolkit for agent-based inference
Inference serving is increasingly open source. But inference optimization still happens inside each provider, and the optimized deployment remains tied to its proprietary stack. This has to change.
Introducing Presets: an open-source toolkit for agent-based inference optimization and a portable format for the result.
Deploying optimized Kimi K3 to any cloud, Kubernetes cluster, or bare-metal fleet, whether on NVIDIA, AMD, or other silicon, should be as simple as deploying a Docker image.
dstack.ai/blog/presets/
Here's one example of using the presets toolkit to optimize Qwen3.8-27B on a single @AMD MI300X.
Through compound learning across linked sessions and source-level patches, performance improved from 311 to 495 tok/s: +59%.
The resulting preset can be deployed on any AMD cloud,
i've been saying for a while now that we need a good harness for tuning inference, but also deploying it.
@dstackai has picked up the ball and done it. glad to partner with them for years now. keep up the great work.
Inference serving is increasingly open source. But inference optimization still happens inside each provider, and the optimized deployment remains tied to its proprietary stack. This has to change.
Introducing Presets: an open-source toolkit for agent-based inference
Node groups are built for workloads like RL training: training jobs, inference, and sandboxes can now run side by side in one run, each on its own hardware.
Here's an example that uses groups to run @raydistributed on dstack, fine-tuning an agent with RAGEN and @verl_project: dstack.ai/docs/examples/…
dstack 0.21.2 is out 🚀
Distributed tasks now support node groups: one run can mix roles and hardware. Each group defines its own nodes, resources, commands, and ports, e.g. a CPU head node driving GPU workers in a single task.
Also in this release: the @AMD Developer Cloud VM image moves to ROCm 7.14, plus a major bugfix.
Full release notes 👇
github.com/dstackai/dstac…
dstack 0.21.1 is out!
Presets: the inference optimization agent can now patch source code (serving framework, kernels, etc). Also, sessions can be linked, so each new one starts from the previous best and pushes further.
Gateways: replicas are now fault tolerant and support in-place scaling.
Plus lots of other improvements: github.com/dstackai/dstac…
Super excited about this release. Most of the work went into presets.
If you spend time optimizing inference, this is for you. You point the agent at a model and it figures out how to serve it best, down to patching kernels when flags aren't enough. A built preset can be
Super excited about this release. Most of the work went into presets.
If you spend time optimizing inference, this is for you. You point the agent at a model and it figures out how to serve it best, down to patching kernels when flags aren't enough. A built preset can be deployed anywhere.
Works on @nvidia, @AMD, and @tenstorrent, any cloud, K8S, or bare metal.
dstack.ai/docs/concepts/…
A longer write-up with results from our own sessions is coming soon.
dstack 0.21.1 is out!
Presets: the inference optimization agent can now patch source code (serving framework, kernels, etc). Also, sessions can be linked, so each new one starts from the previous best and pushes further.
Gateways: replicas are now fault tolerant and support
Find the updated docs on presets at dstack.ai/docs/concepts/…. Can't wait to see what you build with them! Any feedback is very welcome.
A long write-up on how presets find the optimized baseline and streamline the discovery of new kernels is coming soon!
Autoresearch is a real shift in how AI itself is built. @transformerlab just launched Primus, a service that automates the full research loop.
It reads papers, forms hypotheses, writes code, runs experiments, and writes the final paper.
Also great to see @dstackai helping with compute orchestration!
I am deeply proud to finally share what we’ve been building. We're calling it Primus. With a single prompt, it does the full job of a Machine Learning researcher from prompt to paper.
Our goal is massive: we believe this can change the trajectory of how all scientific research
2K Followers 262 FollowingYour private way to access all of the latest and greatest AI models and tools via Bitcoin, Crypto, and credit card top-ups. Pay-per-use, no subscriptions!
545 Followers 2K Followingpowershell and llm inference struggles & shitposts |
local AI open source must win |
SRE @Proofpoint (opinions my own) |
left wing working class voter
636 Followers 7K FollowingBuilding @ https://t.co/5dUr2SjUOu
They are not our dogs
I don’t want to live the wrong live and then die.
📺The Leftovers, Station 11, Succession
1K Followers 3K FollowingShaping Tomorrow in AI Infra & Energy | VP Innovation @TensorWave | Cohost @LiveEIR & Powering the AI Stack | Redistributing the future in KC
6 Followers 730 FollowingStop letting a piece of Javascript bully your best customers behind your back. We let your backend talk like a normal human being who actually likes people."
1K Followers 3K FollowingAgent @yipitdata. RL Infra and pretrain research at night, sometimes photography and poetry. writer account @orange97648. Views are my own
1K Followers 109 FollowingHi, this is https://t.co/EO7MXLjjSU official account. We build systems for efficient LLM serving, including KTransformers, Mooncake and AgentENV.
18K Followers 204 FollowingLarge Model Systems Organization: We developed SGLang @sgl_project (https://t.co/OjwQadINKU), Chatbot Arena (now @arena), and Vicuna!
15K Followers 4K Followingdo things for AI privacy :
Trust Machines @PhalaNetwork /
Privacy LLM @redpill_gpt /
🦞 icloud for AI agent: https://t.co/4MpkGHHOhb /
549 Followers 2K Following📈 AI Engineer @ https://t.co/Du2lQ9AFgU 🗞 Ho scritto spiegoni @ilpost 🎓 MSc Econ & Stats @LaStatale 🎓 BA Filosofia @UniBergamo & @SorbonneParis1
59K Followers 115 FollowingAdvancing AI innovation together. Built with devs, for devs. Supported through an open ecosystem. Powered by AMD.
#TogetherWeAdvance
2K Followers 1K FollowingEngineering leadership @LambdaAPI, open source, dev tools. Computers were a mistake. I don’t speak for my employer. Blog: https://t.co/hqx5umkAO5