Provisioning
We stand up your cluster from a tested, repeatable build: control plane, GPU compute, storage, and networking. Nodes go through burn-in before they run a single job, so bad GPUs and NICs turn up before your users find them.
Provisioning · Kubernetes & Slinky/Slurm · HPC help desk
We deploy and manage Kubernetes and Slinky clusters (Slurm on Kubernetes) using hardware you own or a cloud account you already have. You keep the metal, the accounts, and the data. We do the architecture and run the system.
Your researchers get a production HPC cluster with the Slurm interface they already know, built on Kubernetes and running on infrastructure you own or rent directly. We do the build and the day-to-day operations, and we answer your users' tickets. We are your cluster team.
We are a small shop, and we design each cluster around the workloads it has to run, the hardware and network you already have, and the people who will use it. Then we build it as code so it stays that way.
A cluster that's stood up correctly, with Kubernetes and Slinky/Slurm kept running, and a help desk backing your users.
We stand up your cluster from a tested, repeatable build: control plane, GPU compute, storage, and networking. Nodes go through burn-in before they run a single job, so bad GPUs and NICs turn up before your users find them.
We keep Kubernetes healthy. Upgrades, patching, scaling nodes in and out, networking, storage, and monitoring all run through GitOps, so the cluster's state stays versioned and reproducible instead of accumulating hand-edits.
We run the scheduler layer: partitions, QoS, preemption, accounting, and the Slinky operators behind Slurm on Kubernetes. Jobs land where they should, and usage is tracked.
Direct support for your users and researchers on job submission, partitions and QoS, containers, software stacks, and workflow tuning. The engineers answering tickets run HPC systems for a living.
A single-site reference layout: what each tier is physically made of, how many of them there are, and which networks each one attaches to.
Bare metal, cloud, rented GPU capacity, or a mix. The same reference architecture, adapted to where your compute lives.
We work with cloud providers and colocation facilities regularly, and we know what their quotes tend to leave out. If you are still deciding where to put the cluster, we can help you collect and compare bids from several of them, then size the architecture around what you actually buy.
The same holds for the systems you already run. Plenty of organizations have an identity provider they are not going to give up, and building the integration between it and the cluster's identity and account layer is part of the boutique work. We develop that integration with you, so your people reach the cluster with the accounts they already have.
On-prem racks, a colocation cage, or a lab room in the basement. We handle OS images, network and fabric bring-up, and node lifecycle with Packer, MAAS, Warewulf, and Terraform. A rebuilt node comes back identical to the one it replaced.
GPU instances from the cloud provider you already buy from, inside your project and under your billing and security controls. We bring the reference architecture. You keep the account, the commitments, and the data.
Bare-metal GPU capacity rented from a specialty provider works the same way. Give us nodes and a network, and we will build the cluster on top of them.
Keep a steady-state cluster on your own metal and burst to cloud when a deadline hits. One Slurm interface, consistent images and software stacks, and accounting that covers both.
A modern, GitOps-native stack, configured, tested, and operated by practitioners. These are the parts that have to come together, and how they stack up.
Every box below is defined here, reviewed as a pull request, and applied as code.
Runs the cluster and the scheduler. No user workloads land here.
RKE2
Argo CD
SchedMD slurm-operator
Slurm controller
Slurm accounting on MariaDB
Slurm places Kubernetes workloads
Cilium, kube-proxy replacement
External Secrets Operator, cert-manager
The shared services behind every job and every login.
Slinky loginset
Authentik, SCIM provisioning
nsscache
NFS on ZFS, BeeGFS
csi-driver-nfs, local-path-provisioner
Prometheus, Grafana, Alertmanager
Loki, Alloy
tailnet access to UIs and the API
Where jobs run. Identical from one node to the next.
Slinky nodeset
NVIDIA and AMD
DCGM, AMD device metrics exporter
node-exporter
Enroot, Pyxis
RDMA over InfiniBand or Ethernet
NFS and BeeGFS clients
tailscaled, Tailscale SSH
Your researchers get the Slurm they already know, with srun, sbatch, and sinfo running as Kubernetes workloads underneath. Familiar HPC ergonomics on modern orchestration.
RKE2 with a high-performance CNI, topology-aware scheduling, and device plugins configured for GPU and fabric-heavy workloads.
NVIDIA CUDA or AMD ROCm drivers, container toolkits, and RDMA fabric, configured and validated. Enroot and Pyxis for container jobs out of the box.
Shared home and scratch storage, plus centralized user identity and POSIX accounts that keep working when the identity provider hiccups.
Prometheus, Grafana, and GPU health metrics, plus Slurm accounting backed by an in-cluster database. Usage and utilization stay visible.
Add Tailscale for identity-based SSH and admin access without public IPs or firewall wrangling. Optional, and easy to layer on later.
Kubernetes workloads placed by Slurm instead of by a second scheduler. Containerized services and ordinary sbatch jobs draw from one pool of nodes under one set of priorities, QOS limits, and fairshare.
Launch a vLLM server, a Jupyter notebook, or an isolated VM from a button. Each lands as a Slurm job under your accounting and preemption rules, so an ordinary sbatch can take the node back.
Nothing about your cluster is configured by hand. Everything is defined as code and changed through reviewed pull requests, split across a platform repository we maintain and an apps repository you own.
The platform repository holds the reference architecture and the integration work that makes the pieces fit: Kubernetes and Slinky configuration, storage, identity, monitoring, and the tooling that glues them together. We maintain it and it stays with us. That is deliberate, and it is what you are buying. A fix we make on another cluster this month lands on yours, because you are getting a product we keep sharpening rather than a one-off build that starts aging the day it ships.
The apps repository is yours. It comes in as a submodule, and it holds the workloads your users actually run. A vLLM server behind an OpenAI-compatible endpoint, a Jupyter notebook, and an isolated VM are already built and tested, so putting those on your cluster is configuration rather than a project. Anything specific to you that deploys through slurm-bridge, we build with you in your repository, and you keep it.
Both sides work the same way. Changes are proposed, reviewed, and applied as code, so the cluster's state is versioned, auditable, and reproducible rather than hidden in someone's terminal history. Open a pull request against your apps and work through it with us.
Your admins and researchers get skills that answer cluster questions in plain language, wired to the same read-only interfaces we use.
Your cluster ships with agent skills your team runs from Claude Code or another MCP-capable client. They reach the cluster through read-only tools over Slurm, Prometheus, Loki, and the Kubernetes API. Ask what is idle, what is blocking the queue, who is over quota, or why last night's run failed, and the answer is built from the cluster's own telemetry instead of guessed.
Read-only is the default, and it is the point. A skill that finds a quota problem writes the fix as a diff against your policy file for someone to review and merge. Nothing an agent concludes reaches the live cluster except through the same pull request every other change goes through.
We write skills for your workflows the same way we integrate apps. If your team keeps answering the same question by hand, it should be a skill.
Take the whole thing or just the parts you're short on. Engagements move between these as your team grows.
A defined build. We design the architecture, provision the cluster on your infrastructure, validate it with real workloads, and hand over a documented, running system with your team trained to operate it.
Your admins keep the keys and the day-to-day work. We cover the deep end: upgrades, scheduler tuning, incident response, and the changes nobody on staff has done before.
We operate the cluster and staff the help desk. You keep the hardware, the accounts, and your apps, and you keep a support relationship instead of a hiring problem.
Simple, cluster-size-based pricing for our services: provisioning, Kubernetes and Slinky/Slurm management, and the HPC help desk. Tell us your GPU count and where the nodes live, and we'll size a quote.
Hardware, colocation, and cloud compute are purchased on your own accounts and contracts. We don't resell capacity or mark it up. You pay your provider directly, and you pay us for the engineering, including the work of sourcing and comparing quotes on your behalf.