dstack - Unified GPU Orchestration Across Clouds and Hardware
dstack is an open-source control plane for provisioning and orchestrating GPU workloads across NVIDIA, AMD, TPU, and Tenstorrent on any cloud, Kubernetes, or bare metal.
Tag
14 posts tagged #gpu
Browse 14 posts tagged GPU, including practical setup notes, reviews, comparisons, and workflow patterns for engineers working with AI tools.
dstack is an open-source control plane for provisioning and orchestrating GPU workloads across NVIDIA, AMD, TPU, and Tenstorrent on any cloud, Kubernetes, or bare metal.
A fast CLI and MCP server for managing Lambda cloud GPU instances via natural language commands in Claude Code or other AI assistants.
Velda turns any dev environment command into a cloud batch job. No Dockerfiles, no image registries, no Kubernetes manifests — just prefix with vrun and scale.
Access H100, H200, B200 GPUs on demand from 30+ cloud providers via a single API. Compare prices, deploy clusters up to 64 GPUs, no long-term commitment.
Chamber is an AIOps platform that puts an AI agent between your team and your GPU fleet. Ask in plain language, get root-cause analysis and automated remediation across Kubernetes and Slurm.
Unofficial CLI and MCP server for managing Lambda cloud GPU instances from the terminal or inside AI editors like Cursor and VS Code.
Chamber deploys an AI agent named Chambie that monitors, root-causes, and fixes GPU issues across AWS, GCP, Azure, and on-prem clusters — answering in Slack. YC W26.
Deploy APIs, AI inference, and databases on Koyeb's global serverless infrastructure. GPU instances, per-second billing, autoscaling to zero. CLI setup guide.
YC W26 inference API with custom IonAttention engine. Drop-in OpenAI replacement, built for NVIDIA Grace Hopper. Per-second billing, no cold starts.
Shadeform aggregates 8+ GPU providers into a single marketplace and API so you can find, reserve, and launch A100s and H100s without hyperscaler lock-in.
High-throughput GPU storage engine built on Rust and io_uring — GPUDirect Storage, RDMA, erasure coding, 6x higher throughput than conventional object stores.
Expanse reads your job's source code, submission script, and hardware telemetry to predict actual GPU needs—cutting 59% datacenter waste before jobs run.
AI agents that autonomously monitor, root-cause, and remediate GPU infrastructure issues. Reduce compute costs, improve GPU utilization, and accelerate ML research. YC W26.
Chamber is a YC W26 AI teammate that autonomously monitors, root-causes, and remediates GPU infrastructure issues across multi-cloud.