AI & Agents · Systems

strata-swift-flash-next-rtx3090

Runs Swift 1.5 Flash Next, a 125B-class mixture-of-experts model, on the single RTX 3090 in Zeus. Expert weights stay in 192 GB of host RAM with a selected-expert GPU cache, MTP speculative decoding and a 262,144-token context. It is the machine's default inference backend, measured from 4K to 262K prompts and managed as a Nix-declared service behind Open WebUI.

Active · 2026

Strata · CUDA · Nix · RTX 3090

strata-swift-flash-next-rtx3090

My previous build runs a 27B Swift entirely on an RTX 3090. I spent most of that project rebuilding its quantised heads and speculative drafter to get more tokens out of 24 GB of . The model was small enough that I could keep its weights on the card.

Swift 1.5 Flash Next is a different problem. It is a 125B-class mixture-of-experts model, and the IQ3_S checkpoint I wanted comes in two GGUF shards totalling about 84 GB. I have one RTX 3090, but Zeus now has 192 GB of system RAM. Strata can run the model with most of its expert weights in RAM and selected experts cached on the GPU.

I wanted to run this as my everyday local inference backend, not as a demo that answers one prompt. It needed to work with my coding agents, accept tool calls, survive service restarts and handle long conversations. I also wanted to know how long a large prompt actually takes to read. A nominal 262,144- context is only useful if the model can retrieve information from it.

The deployment is now my default on Zeus. Strata serves Swift 1.5 Flash Next IQ3_S, and ai.liamwh.com uses it through Open WebUI. That host sits on my tailnet, so the link only resolves for me. The path there was less about model conversion than CUDA linkage, systemd unit behaviour and several tests that initially measured the wrong thing.

The machine and the launch profile

Zeus runs Ubuntu 24.04 on a Ryzen 9 9950X, with four 48 GB DDR5 DIMMs, one 24 GB RTX 3090 and an NVMe SSD. The RAM is configured at DDR5-4800. It passed a short stress test at that speed, but I have not yet done a bootable overnight memory test.

The Flash Next checkpoint is too large for VRAM. Strata loads expert weights into system memory and caches some of them on the GPU. An expert already in the GPU cache can run there. An uncached expert costs more because the engine has to fetch its weights over PCIe or use the CPU path. Which experts remain on the GPU matters, and the hit rate changes with the prompt.

One successful load reported 46.84 GiB of experts resident in host memory and 7,465 cached experts using about 14.15 GiB of VRAM. That is the engine’s allocation report, not the checkpoint’s total file size or a constant for every run.

Here is the configuration I ended up running.

SettingDeployed value
EngineStrata 0.1.41, provisioned binary with a checked SHA-256
ModelSwift 1.5 Flash Next GSQ-RCO, IQ3_S
Context262,144 tokens
KV cacheint8, with KV streaming
Resident KV window32,768 tokens
SpeculationMTP, four draft tokens per verification window
GPU expert placementProfile-guided automatic cache
VRAM reserve2,048 MiB
Client APIOpenAI-compatible and Anthropic-compatible endpoints, authenticated

The 262,144-token limit is configured and exercised below. It does not mean a 262k prompt is cheap, or that the model is equally reliable at every position in a long document.

Getting the weights onto Zeus

I used UkisAI’s Swift 1.5 IQ3_S release rather than a smaller . The two GGUF shards total roughly 83.74 GB. I pinned the Hugging Face revision and checked both shards against the publisher’s SHA-256 manifest. I also fetched and verified the vision projection asset, although my first deployment and the tests in this article concern text inference. A verified vision file is not proof that image input works end to end.

The draft layer also needed conversion into Strata’s runtime layout. I kept the original model files, prepared draft and expert assets, and engine binary under ~/models/strata. None of the large model files belong in the Nix store.

The first working service came from the upstream checkout with a local Python environment. It proved that this model could load and generate on Zeus. It was not the setup I wanted to maintain. My workstation already uses Nix, Home Manager, systemd user units and for the rest of the local LLM stack.

What Nix manages, and what it doesn’t

Zeus is Ubuntu, not NixOS. Home Manager generates the user units, scripts, launch parameters and most runtime dependencies. The API key comes from an existing SOPS-managed local LLM secret. The backend listens on the existing inference port, 8090, so my clients do not need a new network address whenever I change engines.

I initially tried to build the entire Strata engine under Nix. The first build compiled, but its Python server could not import strata_tokenizer. After I added the module path, it failed on a missing regex dependency. A C++ binary reporting its version had told me nothing about whether the HTTP server could start.

A cheap package smoke test now checks the Python imports and API behaviour without loading the full model. It exercises configuration and authentication, including a rejected request without a key. Those errors should appear before a service tries to load tens of gigabytes into memory.

CUDA was harder. The Nix-built executable initially picked up a driver library from /run/opengl-driver/lib instead of the real host libcuda.so.1. The resulting error was misleading.

CUDA driver version is insufficient for CUDA runtime version

Running the executable outside the systemd sandbox failed in the same way, which ruled out systemd permissions as the immediate cause. Preloading the host driver library fixed that part. The unit also needed a writable, persistent CUDA_CACHE_PATH and the matching CUDA toolkit’s bin directory on PATH. cuBLAS invokes ptxas during some JIT compilation, so an executable that starts successfully can still fail on its first substantive prompt.

There were two more problems. A re-shipped NVIDIA cuBLAS 13.1.1.3 archive broke the CUDA build through the version of nixpkgs I was using. I added a separately pinned nixpkgs-strata input at revision 8edc0c72 to keep the known-working CUDA runtime. Separately, engines compiled inside the Nix sandbox failed cublasGemmEx on prompts of at least roughly 1,000 tokens. Different compiler and optimisation combinations did not fix it, while an engine built through Strata’s upstream host-toolchain flow worked.

I stopped trying to force the native engine into the sandbox. The deployed engine is a host-built binary provisioned under ~/models/strata/engine/strata. Its expected SHA-256 lives in version control and the service checks it when starting. The Python serving layer, launch scripts and service configuration remain Nix-managed.

That is a compromise. The deployment is declarative about which engine artefact may run, but the engine itself is not reproducibly compiled inside the Nix sandbox. I prefer to state the difference rather than call the entire executable a Nix build.

Why Home Manager kept replacing Strata

I already had Swift-Qwen, syv-vLLM, llama.cpp and ComfyUI sharing the RTX 3090. The inference services have conflicts so that starting one stops the others. Strata joined that group and uses port 8090 when active.

The first stable Strata start lasted only seconds. Then Swift-Qwen started and Strata stopped. The timing initially implicated Stile because it manages secrets for several backends and the root-owned registry on Zeus did not yet recognise Strata. Reading its Rust source changed the diagnosis. When no declared backend is active, stile reconcile returns an error. It does not start Swift-Qwen.

I reproduced the switch during home-manager switch and traced the systemd start requests. Home Manager’s sd-switch started services that were enabled but inactive. Swift-Qwen was still enabled as my boot default. Every switch could therefore restore Swift-Qwen, which stopped Strata through Conflicts=.

I changed the Home Manager configuration to:

systemd.user.startServices = "suggest";

That means Home Manager no longer starts enabled-but-inactive user services during activation. After the change, a switch returned successfully with Strata still active, Swift-Qwen inactive and the model still loaded. The trade-off is that I must explicitly start new services after deploying them; their normal boot enablement remains separate.

Stile needed a smaller correction. Its repository registry contains the Strata entry, but the deployed registry on Zeus does not list Strata yet, so reconciliation for it stays a manual operator task. Until then, local-llm-sync avoids asking Stile to reconcile a backend it does not know.

The test failures I caused

The service initially ran an inference test as part of systemd startup. I told the model to answer STRATA_OK, then searched the raw streamed response for that literal string. The engine generated the expected answer, but the check failed.

The problem was server-sent events. Strata sent the text in separate JSON deltas, including STR, ATA and _OK. The complete string never appeared as contiguous bytes in the HTTP response. My first attempt at fixing the test concatenated JSON metadata along with the content, which was no better.

I eventually parsed the SSE data: events as JSON, joined the actual text deltas and checked the completed stream. I also moved full inference checks out of the startup gate. Readiness now tests whether the process, loaded model and authenticated API are available. Tool calls and answer content belong in a separate acceptance suite.

The bad test had an expensive side effect. A failed systemd start triggered another large model load while processes from earlier attempts were still winding down. Some of those overlapping engines did hit real cuBLAS allocation failures. I spent time changing the VRAM reserve before separating resource contention from the faulty output assertion. The deployed reserve went back to 2,048 MiB. I do not have a controlled reserve-size sweep that would justify calling that value optimal.

The long-context benchmark had two mistakes of its own. With reasoning_effort set to none, the model missed the retrieval test even at 4K. After correcting that, repeated identical prompts sometimes received a cached earlier answer. One apparent 51,000-token-per-second prefill was a cache hit, not a record-breaking GPU. I added a unique value to each test prompt so each run performed the work I intended to measure.

What I measured

I tested retrieval of a specific marker embedded in generated documents at five context sizes. These are the corrected runs, with the normal reasoning setting and per-run cache busting. The token counts are the actual prompt sizes, not the context limits.

Configured contextInput tokensTime to first tokenPrefill tokens/sDecode tokens/sRetrieved marker
4K3,8381.84 s2,08593.8Yes
32K30,46611.86 s2,568105.3Yes
64K61,01324.10 s2,53296.6Yes
128K123,05550.34 s2,44495.2Yes
262K247,146107.67 s2,29571.6Yes
Paired bar charts of the Strata context ladder on one RTX 3090. Left, time to first token grows from 1.84 seconds at the 4K rung to 107.67 seconds at the 262K rung, with 11.86, 24.10 and 50.34 seconds in between. Right, decode speed stays between 93.8 and 105.3 tokens per second for the first four rungs and drops to 71.6 at the 262K rung. The marker was retrieved at every rung.
The two costs of a long prompt. Reading it grows with size, while writing the answer stays near 100 tokens per second until the largest rung. Actual prompt sizes ranged from 3,838 to 247,146 tokens.

The longest prompt was roughly 247k tokens and the marker was returned correctly. That establishes that a real request near the configured 262,144-token limit worked. It does not establish broad reasoning quality at that length. This was a marker-retrieval test, not a suite of repository-wide coding tasks with 247k-token histories.

The other immediate result is the cost of a very long first prompt. The 247k request took almost 108 seconds to reach the first token. speed was still about 72 tokens per second once generation began, but the prompt-processing time would be noticeable in an interactive coding session. That matters more to me than the context number alone.

Shorter tests give different speeds. One short task decoded at about 75 tokens per second. Some coding prompts ran at roughly 88 to 100 tokens per second. A quote workload with an 8,259-token input decoded at about 100 tokens per second and recorded a match score of 0.991. These are results from the recorded Strata workloads, not medians over many independent boots or proof that the model is better at writing code than my old 27B deployment.

I deliberately did not benchmark the 27B service in this project. That comparison would need matched coding tasks and a separate account of model quality, total reasoning tokens and completion rates. It would not answer the immediate question of whether I had a reliable Flash Next server.

The configuration uses four speculative MTP tokens, a 32k resident KV window and profile-guided expert caching. I have not published a controlled ablation of speculation on versus off, KV precisions or different expert-cache budgets. The recorded speeds describe this working configuration, not the fastest configuration proven possible on Zeus.

Testing the API rather than assuming compatibility

Strata exposes OpenAI-compatible and Anthropic-compatible endpoints. I ran the acceptance suite against the deployed model, not the lightweight HTTP smoke-test server. It checked that an unauthenticated model request returns 401, a streamed chat response completes with valid SSE events, and a forced function call returns parseable arguments. It also exercised reasoning-effort controls, /v1/responses, /v1/messages, cancelling a stream and two requests queuing behind each other. The full suite passed.

Two service restarts passed, and a Home Manager activation left Strata running. I also switched between Strata and Swift-Qwen and back. Strata’s /metrics endpoint returned 101 Prometheus series under bearer authentication, using names compatible with the vLLM monitoring conventions already on Zeus. That doesn’t mean every external Prometheus scrape has been configured; the optional fleet scrape still needs a credential on the collecting host.

A short run of eight sequential generations completed without errors. I do not treat that as a long-duration stress test. About 15 GB of swap was already allocated on the workstation during the final inspection, although the short inference run did not show a slowdown consistent with active swapping. More extensive RAM testing remains worthwhile, particularly with four recently installed DIMMs.

Making it the default for Open WebUI

Until this project, Open WebUI was effectively configured for llama.cpp. It ran as a rootless container, pointed at port 8090 and passed the llama.cpp API key. When Swift-Qwen or Strata owned that port, the web app could be up while its model connection failed authentication.

I kept the existing ai.liamwh.com site, its database and its container image, pinned at a specific Open WebUI release. The wrapper now checks which backend owns the local inference port. Strata, Swift-Qwen and syv-vLLM share the existing local LLM credential and use the regular OpenAI provider. The llama.cpp fallback retains its own key and provider configuration.

Open WebUI stores some of these settings in SQLite, and the saved values override environment defaults. Changing a container environment variable alone was not enough. The wrapper now backs up webui.db before changing relevant connection settings, then seeds the API connection and default model for the selected backend. It uses the actual Strata API model ID, strata, and gives that entry the display name “Swift 1.5 Flash Next (Strata)“. The model description states the 262,144-token context without setting that number as a default completion length.

I also stopped putting the API key in the visible podman run command line. Podman receives it through a permissions-restricted environment file on the user’s runtime filesystem. Open WebUI still needs the credential in its runtime configuration database, so the database and its backups remain sensitive files.

Strata is now enabled as the boot-default inference service, and Swift-Qwen is no longer enabled at boot. The other backends remain installed for manual switching. Open WebUI orders its startup after the inference units but does not require them or pull them in. That distinction matters. If restarting the web UI also started Strata, it would take the GPU back from a backend I had deliberately selected.

There was one more systemd failure in this change. I initially restarted Open WebUI directly from an inference service’s ExecStartPost. Open WebUI had an After= relationship on that service. The restart waited for inference startup to finish, while inference startup waited for the restart. It hung until the systemd timeout. I changed backend-switch reconciliation to schedule the web UI restart through a separate transient systemd unit. The wrapper compares the backend with its last-seeded state, so a switch updates the web UI once, but a same-backend Strata restart leaves it alone.

I checked the result in an authenticated browser session against ai.liamwh.com. A new chat selected “Swift 1.5 Flash Next (Strata)” by default, and a real streamed response displayed through the web interface, including its reasoning indicator. The nine pre-existing conversations were still present. I then switched to Swift-Qwen, verified the default changed to qwen3.8-27b, and switched back to Strata. The web UI remained accessible during the switches.

The final local check reported Strata active, Swift-Qwen and syv-vLLM inactive, Open WebUI active, the default model set to strata, and the HTTPS ingress responding with status 200. I have not rebooted Zeus to test the new boot default from a cold machine. The enablement, ordering and controlled restart paths have been checked without a reboot.

What remains

The project now has a working 262k-context Strata service, API tests, long-prompt measurements, Home Manager integration and a usable Open WebUI front end. The core infrastructure changes, including the Open WebUI promotion, are committed in my private infrastructure repository, and the Strata providers for my coding agent are committed in my dotfiles. Both repositories are private, so I don’t publish their revisions here.

I still need to run the root-owned Stile registry bootstrap to make Strata a fully registered broker backend, provision the optional external Prometheus scrape credential, and perform an offline overnight memory test. An actual reboot test is also outstanding. More repeated, controlled benchmarks would be needed to choose between alternative Strata tuning settings or make a claim about coding-task success against another model.

For now I use the configuration above. Its limits are measured closely enough for me to know what I am getting: around 100 decode tokens per second on some coding prompts, and a working 247k-token retrieval prompt that takes nearly two minutes before the first token. The rest of the work is about reliability over time and whether those numbers translate into completed coding tasks.

Open chat

Interested in working together? Reach out.

Strategy, architecture, and implementation — from workflow to production.

© 2026 Liam Woodleigh. All rights reserved.