Systems · AI & Agents
Zeus
My Linux development workstation, which also serves local LLMs to my coding agents through vLLM and runs image generation. It shares one RTX 3090 between the two workloads, and Prometheus, Tempo and Loki feed Grafana with the telemetry to see what limits them.
Active · 2026
vLLM · CUDA · Nix · Grafana

Zeus is the machine I write software on, and it is also where my coding Coding agent An AI model given tools to edit files and run commands, so it can carry out a programming task end to end instead of only suggesting code. get their model. It runs Ubuntu 24.04 with a Niri Wayland session for day-to-day development, and vLLM An open-source server for running large language models (LLMs) on GPUs. The v comes from the virtual-memory technique it uses to manage the context. It handles memory, batching and speculative decoding. serves local LLM Large language model. A model trained on a lot of text to predict the next token, which is what chat assistants run on. on the same box behind an OpenAI-compatible API. When I need images rather than text, ComfyUI takes over the GPU.
I built it to answer questions I couldn’t answer from a hosted API. How practical is good local inference on consumer hardware? Can one machine be a pleasant workstation and a model server at the same time? When throughput disappoints, what is actually the limit: compute, memory bandwidth, the KV cache, batching, the CPU or scheduling? And can a local model serve an interactive coding agent well enough that I stop reaching for the cloud?
Why run models locally
Hosted models are convenient, and for plenty of work they are the right choice. Running locally buys something different: prompts and code never leave the machine, I choose the model, the Quantisation Storing a model's weights with fewer bits, for example 4 instead of 16. The model gets smaller and faster to read from memory, at some cost in accuracy. and the serving flags, and I can see and change every layer between a request and the GPU. The marginal cost of another request is electricity, which makes it cheap to run experiments I wouldn’t pay per token for. None of that makes local inference cheaper in general. A GPU sitting idle costs money too.
The hard part is that Zeus is not a dedicated server. It is my daily workstation, so the serving stack has to fit around an editor, a browser, builds and tests, without making any of them worse.
Architecture
The machine is an AMD Ryzen 9 9950X with 96 GB of DDR5, 4 TB of NVMe storage and one NVIDIA RTX 3090 with 24 GB of VRAM Video RAM, the GPU's own memory. The RTX 3090 has 24 GB, and the weights, the context and working buffers all have to fit in it.. I sized the power supply and cooling for sustained inference rather than bursty desktop use, and left room for a second GPU.
There are three layers. Clients speak the OpenAI API, so any tool that can point at a base URL can use a local model without knowing it is local. vLLM owns model serving. A small control layer decides which GPU-heavy workload is running. Alongside them, an observability stack records what the serving layer and the hardware are doing.
Serving local models
vLLM serves Qwen-family models, including a 27B-class model that is my default for coding work. The OpenAI-compatible endpoint is what makes that practical: coding agents, editor integrations and scripts all talk to Zeus the same way they would talk to a hosted provider, and swapping models is a server-side change.
Coding agents are a demanding workload for a single consumer GPU. They send long prompts full of files and tool output, they reuse most of that prompt on every turn, and they need a large Context window How many tokens the model can see at once, including the whole conversation so far. 114k tokens is a few hundred pages of text. for long sessions. That shapes most of the serving choices: how much VRAM goes to weights and how much to the KV cache, whether prefix caching pays off, and which quantisation keeps quality while leaving room for context. swift-qwen3.8-rtx3090 is one piece of work that came out of this: rebuilding the speculative-decoding parts of a 27B model so it runs faster on this GPU.
Sharing the GPU
LLM serving and image generation both want most of the 24 GB. vLLM reserves a fixed share of VRAM up front for weights and KV cache, and an image pipeline has its own large working set. Running both at once would mean shrinking one of them until neither is useful.
So I don’t try. Zeus treats the GPU as a resource held by one heavy workload at a time, and I built tooling that switches between them: stop the current workload and start the other. The switch is an explicit operation rather than something that happens when two processes collide and one of them fails to allocate.
Observability
nvidia-smi tells me the GPU is busy. It doesn’t tell me whether that busyness is useful. A GPU can report near-full utilisation while requests queue, the KV cache is full and each new Token A chunk of text the model reads and writes, usually a word or part of one. Speed, context length and output are all counted in tokens. waits on memory bandwidth rather than compute. Utilisation says a kernel was running, not how much work it did.
To see the difference I run the Grafana stack locally. Prometheus scrapes vLLM’s /metrics endpoint alongside NVIDIA GPU telemetry and host metrics. vLLM’s traces go through Grafana Alloy to Tempo, and its logs through Alloy to Loki, so a slow request can be followed from its trace into the logs around it. Grafana is the one place I look at all of it.
The dashboards are organised around relationships rather than single numbers:
- Queue depth and concurrency against KV-cache utilisation. When the cache fills, vLLM can’t admit more requests, and queueing grows even while the GPU looks busy.
- Prompt throughput against output throughput. Prefill and Decode Generating the answer one token at a time after the prompt has been read. Decode speed is measured in tokens per second. stress the GPU differently. Agent workloads are prefill-heavy, and prefix-cache hits change that balance.
- Time to first token against end-to-end latency. TTFT tracks queueing and prefill. Time per output token (inter-token latency) tracks decode. Which one dominates end-to-end latency says where to look.
- GPU utilisation, VRAM, temperature and power against all of the above. These show whether the hardware is the constraint or just the thing that is visibly busy.
- CPU, memory and host health. Zeus is still a workstation, so I need to know when inference is slowing down everything else, and the reverse.
The measurements I take for future write-ups come from these dashboards.
A workstation first
Zeus is still my primary development machine. My user environment is managed with Nix Home Manager on top of Ubuntu, so shells, editors, CLI tools and their configuration are declared rather than accumulated, and I can rebuild the environment instead of repairing it. The inference stack has to live alongside that without taking it over. A model server that makes the editor stutter would defeat the point of the project.
What’s next
None of this is done yet:
- A second GPU. I plan to add another high-VRAM NVIDIA card, and the platform was built with that in mind.
- Comparative benchmarks. Single GPU against two, and tensor-parallel serving of one model against independent workers each serving their own. I also want to compare model sizes and quantisation strategies, and measure concurrency, long-context behaviour, TTFT, throughput and power efficiency on real workloads.
- Better workload scheduling. Moving from manual switching between LLM serving and image generation towards something that decides for me.
- More local coding-agent work. Finding where local models hold up in real agent sessions and where they don’t.
I’ll publish the results as write-ups with the measurements behind them.