claudeers.
// Other

slotstream

Run a 105 GB AI model on a Mac that can't hold it. Slotstream streams Qwen3.8-Flash-Next (125B mixture of experts) from your SSD and caches the busiest exper…

// Other[ cli ][ api ][ desktop ][ web ][ claude ]#claude#apple-silicon#claude-code#inference-engine#llm#llm-inference#local-ai#local-llm#other◷ MIT$open-sourceupdated 8 days ago

Install with your AI

Paste into Claude Code, Cursor, or any agent — it reads the repo and wires the tool into your project.

Install and set up slotstream (release-binary project) into my current project.
Found on https://claudeers.com/slotstream
Repo: https://github.com/carloslfu/slotstream
Homepage/docs: https://sevrahq.com
Detected install method: release-binary → inspect the README
Category: other. Platforms: cli, api, desktop, web.
Read the repo's README for exact setup and env vars, then install it and wire it into my project.

Claudeers Health Verdict:
unknown; community-verified: false. Confirm the source before running anything.
// or install directly (release-binary)

Grab the latest release asset from GitHub.

# download a build from https://github.com/carloslfu/slotstream/releases
// or clone
git clone https://github.com/carloslfu/slotstream

// compatibility

Platformscli, api, desktop, web
Operating systems—
AI compatibilityclaude
LicenseMIT
Pricingopen-source
LanguageSwift

Get your FREE $2.50 API credits to access TickAtlas financial data ↗

slotstream

Run a 105 GB AI model on a Mac that can't hold it.

Slotstream runs Qwen3.8-Flash-Next, a 125-billion-parameter open model, on Macs with 16 to 64 GB of memory. It keeps most of the model on the SSD and loads the parts it needs as it writes. Our 48 GB M5 Pro measured 15.86 tokens per second at a 22 GB memory target (how it was measured).

Chat with it, ask it about pictures, or code with it: slotstream launch claude starts Claude Code on the local model, and Codex, Pi, opencode and Hermes work the same way. Developers can connect their own apps through its Ollama-, OpenAI- and Anthropic-compatible APIs or its Swift library.

After a one-time download it works offline, with no Python and no cloud account. The whole engine is one native Swift program on Apple's MLX and Metal; see Built native. Every published number has a recorded method, and the experiments that failed stay in the measurements.

Get started · Speed · Guides · Get help

I'm building Sevra on Slotstream: private, personal AI optimized for your computer. Sevra will choose a tested model for your hardware, keep that choice current as models improve, and let you control what it remembers. The Mac app is in development and runs Slotstream in process; see how it is built and join the waitlist. Slotstream's command-line tool, APIs and Swift library remain independently usable.

Who it's for

Slotstream is built for Macs that cannot hold the model in memory: 16 to 64 GB. That is where the engineering, the measurements and the defaults go, so that frontier-class intelligence runs on the Macs most people already own. It also runs on 96 GB and larger Macs, where the model fits in memory, but it is not optimized for them: engines that keep the whole model in memory report faster replies there. See related projects if that is your Mac.

Will it run on my Mac?

You need an Apple Silicon Mac with at least 16 GB of memory, macOS 14 or later, and about 110 GB of free SSD space. Open About This Mac from the Apple menu to check your chip and memory. On an 8 GB Mac even the smallest memory plan doesn't fit, so Slotstream refuses to start instead of swapping. Windows, Linux and Intel Macs are not supported. The hardware guide has the tested macOS versions.

Speed

tok/s means tokens per second; a token is a small piece of text, often part of a word. Reply speeds below describe generation after the model has warmed up.

Our development Mac, a 48 GB M5 Pro, measured 15.86 tok/s with 0.2.19 at a 22 GB memory target, in a controlled benchmark on eight prompts the engine was never tuned on. The engine predicts which experts the next layers will need and reads them from the SSD before they are asked for, which changes speed and never the output. The expert lookahead guide has the measurements behind each release. This historical test used smaller prompt passes and disabled prefix caching, leaving more memory for experts. It is not a measurement of today's automatic configuration. A qualified full-answer baseline on 0.2.23 is still pending.

What to expect by memory

Rough planning ranges for warm replies, from community reports and our own measurements, rounded outward. Faster chips and SSDs sit at the top of each range; other apps and memory pressure pull results down.

Installed RAMEstimated warm reply speedExample automatic context window
8 GBSupport coming soon. The current model doesn't fit yet.Not available yet
16–<24 GB~1–6 tok/s32,768 tokens
24–<48 GB~5–16 tok/s32,768 through 32 GB; 65,536 at 36 GB
48–<96 GB~15–27 tok/s32,768 at 48 GB; 131,072 at 64 GB
96 GB+, the model fits in memory~20–32 tok/s262,144 tokens, the model's full window

Context examples use decimal-GB memory simulations. A Mac's marketed capacity, Metal limits and available memory can produce a different plan; slotstream doctor shows the actual choice.

The middle rows are anchored on our M5 Pro's measurement; the top ends of the last two rows come from a 128 GB M5 Max with a larger, manually chosen memory target, and the last row is outside Slotstream's target range. These are estimates, not limits. The hardware guide has the basis of each range, every result measured on real Macs with credits and test conditions, and every automatic memory plan.

Recent prompt-processing results

The changes shipped in 0.2.23 shorten prompt processing and repeated-history work. These measurements use the same 48 GB M5 Pro at a 10 GB target, with three clean pairs per comparison:

WorkloadMatched controlMedian times, control → enabledMedian paired time reduction
16K inventory prompt, MTP onLarger-read workspace policy off155.22 s → 53.94 s prefill65.40%
2K prose follow-up, MTP offPrefix checkpoints disabled30.73 s → 4.42 s request85.62%

Both comparisons switch a feature off in the same tested binary. They measure prompt processing or a cached follow-up, not an increase in reply tok/s or a whole-release speedup. Times are arm medians; reductions are medians of paired changes. The hardware guide explains the fixtures and the latest audit's exclusions.

A fresh installed-release study used ordinary caching at the same memory target. These are observed first-read ranges and median exact-repeat request times across the prescribed request order, with capped replies:

PromptEligible first reads / repeatsFirst-read prefill rangeMedian repeated request
2K code4 / 313.86–28.11 s2.79 s
2K prose3 / 314.91–26.51 s3.22 s

The desktop load screen passed for the included observations, but several runs had system swap-ins. Request history changed read batching, so these results do not replace the general speed estimates or decode headline. See the measurement and its limits.

Memory and context

Auto mode picks the memory target, cache size, speculative decoding and context window for your Mac. It takes the largest window in the table above that still leaves room for speculative decoding and a complete conversation, given the memory free at startup, without an unmeasured loss of useful expert cache. slotstream doctor shows the choice and why, and --max-context 65536 sets a window yourself, up to 262,144 tokens.

Starting a reply takes time. Slotstream first reads your question and the conversation history, which can take minutes for a long prompt. Follow-up turns reuse unchanged history, a new conversation reuses the system prompt earlier ones started with, and serve --prefix-cache-dir keeps long conversations and shared system prompts on disk so that they survive a restart. The hardware guide has the prompt-reading estimates for each memory size.

Install

Open Terminal and paste this command:

curl -fsSL https://raw.githubusercontent.com/carloslfu/slotstream/main/install.sh | sh

Run the same command to update. If slotstream isn't found afterward, open a new terminal window.

Use it

Check your Mac, then ask for a first reply:

slotstream doctor
slotstream run --prompt "Why is the sky blue?"

doctor checks memory and disk space without loading the model. The first run asks to download it, then prints a reply. This download can take hours, but you only need to do it once. Interrupted downloads resume when you try again. Follow the step-by-step setup for more help.

Guides

What would you like to do?Guide
Chat in Open WebUI or another appConnect a chat app
Ask about a pictureUse an image
Start a coding agent on the model, in one commandUse coding agents
Code with Claude CodeUse Claude Code
Code with CodexUse Codex
Code with Pi or opencodeUse Pi or opencode
Work with files and tools through HermesUse Hermes
Code with fx, Vercel Labs' coding agentUse fx
Fix a problem, move the model, or uninstallTroubleshooting

Install chat apps and agents separately. They provide the interface and tools; Slotstream runs the model. Keep its server running while a connected app uses it. slotstream launch claude (or codex, pi, opencode, hermes) starts that agent already connected, and starts the server in the background first when none is running.

For developers, the engineering guide links to the OpenAI- and Ollama-compatible API references, Swift library, command options, build instructions, and tests. Release notes show what changed.

How it works

Qwen3.8-Flash-Next is a mixture-of-experts model: generating each piece of text uses only a subset of its expert networks. Slotstream keeps shared weights in memory and reads the needed experts from SSD into a cache. Frequently used experts stay in RAM, reducing repeated disk reads.

The whole model stays available even though it doesn't all fit in memory. Slotstream chooses a memory target for your Mac and adjusts its cache as other apps need room. Cache size changes speed without removing experts from the model. The engineering explanation covers the implementation.

Built native

Slotstream is one native Mac program: the command-line tool, the HTTP server, the memory planner, the expert cache and the model itself are Swift on Apple's MLX framework and Metal, with Slotstream's own Metal kernels compiled at run time where a step needed one (the gated-delta recurrence, the selected attention, the expert routing). No Python runtime or interpreter sits between a request and the GPU.

That is deliberate. Built around one model on one kind of hardware, each layer is tuned for the one below it: expert records are read from the SSD straight into the cache slots the GPU computes from, the memory plan is checked against what the process really uses, and a governor resizes the cache while other apps need room. It is also why the engine ships as one file that installs with one command, and why a Mac app such as Sevra can run it in process. The speed on this page comes from measured mechanisms that this control allows, not from the language itself; every published number has a recorded method in the measurements. The trade is that the engine runs only on Apple Silicon; Windows and Linux are planned in Sevra with their own native engines.

Status and limits

  • Not optimized for 96 GB and larger Macs, where the model fits in memory; see Who it's for.
  • One generation at a time: connected apps share the same running model.
  • Conversation length is limited: longer histories take more memory and time. Auto mode picks a window for each memory size, and the coding agent guides include the larger window agents need.
  • No broad benchmarks yet: image input and tool calling have integration tests, but there is no image-accuracy benchmark or completed comparison with other models on the same Mac.

FAQ

Does it work offline?

Yes, after downloading the model. Inference runs on your Mac. Connected agents may still use internet services for web searches or other tools; their settings determine what those tools send.

Why doesn't Slotstream use all of my RAM?

--memory-gb 48 is a maximum process budget, not a promise to keep 48 GB resident. Slotstream allocates the expert cache at load, while conversation state and temporary work grow only when a request needs them. A short request can therefore peak well below the target. The startup report shows the budget, expert cache, runtime allowances and safety headroom separately.

Without an explicit target, Auto uses a 33 GB base ceiling, or 34.6 GB with speculative decoding at the 32,768-token window. This measured default leaves memory for other apps. To use more, stop any running server and preview the plan without loading the model:

slotstream doctor --memory-gb 40

If it fits with headroom, slotstream serve --memory-gb 40 uses that budget with a fixed cache. In the current development version, use --memory-limit-gb instead to choose an upper limit while the cache adapts to other apps. Custom limits can exceed the automatic default; the Mac's supported budget and available memory still bound actual use. Auto keeps expert cache when a larger automatic context would trade it away without a measured benefit. Use --max-context N when you explicitly want a longer window. Leave room for macOS and other apps. See the memory options for details.

Is Slotstream the fastest way to run this model?

On a Mac that cannot hold the model, 16 to 64 GB, it is the way to run it at all, and the engineering goes into making that fast. On 96 GB and larger Macs the model fits in memory and engines that keep it there report faster replies. See Who it's for and related projects.

Why is it written in Swift and not Python?

The engine needs direct control of memory, disk reads and the GPU, and it has to ship as one file that a Mac app can call in process. Swift on MLX and Metal gives that; a Python runtime would put an interpreter and a second process in the way. The Python in the repository is tooling (the reference model the port is checked against, benchmark drivers, release checks), and none of it runs when you use Slotstream. See Built native.

Will this wear out my SSD?

Generation reads the model files without rewriting them. macOS swap adds writes when memory runs short. Automatic memory sizing helps, but a small Mac or an oversized manual setting can still swap heavily. Servers that slotstream launch starts, and serve --prefix-cache-dir, also save each turn of a longer conversation to a cache folder, and the system prompts conversations share, within a disk quota.

Can I run it on Linux or Windows?

Support for AMD and NVIDIA on Windows and Linux is planned for Sevra. It isn't available in the current Slotstream engine.

Can I use a different model?

Not with Slotstream today. Its loader and memory planner are built for this model. See related projects for runtimes with different model and hardware support.

Why this exists

I wanted to run this model on my own Mac, but the standard loader exhausted memory before producing a reply. Slotstream grew out of that experiment. The published measurements include that failed load and the experiments that followed.

The project was also discussed on Hacker News. The questions and hardware reports from that discussion help guide the work.

Support

Report a bug if something doesn't work, or share your Mac's results to help others know what to expect. Reports are credited to their authors. Code and documentation contributions are welcome; see Contributing for the workflow.

Grants and sponsors

Guillermo Rauch's GitHub profile photo
Guillermo Rauch

Slotstream was selected for Guillermo Rauch's personal grants for foundational open-source software. Thank you for supporting its development.

Who made this

I'm Carlos Galarza. I build local AI and make it run efficiently on the computers people already own. Slotstream is the engine, written natively for the Mac and tuned as one system, and Sevra is the private AI app I'm building on it. I also help teams run open models on their own hardware and debug agent workflows. For help or consulting, email me.

Star history

Slotstream GitHub star history, updated weekly

The badge at the top shows the latest star count; this chart is updated weekly.

License

Slotstream is MIT-licensed. The model weights have their own Qwen community license. See credits for the model and code this project builds on.

// faq

What is slotstream?

Run a 105 GB AI model on a Mac that can't hold it. Slotstream streams Qwen3.8-Flash-Next (125B mixture of experts) from your SSD and caches the busiest experts in memory, so it runs on Macs with 16 to 64 GB. One native Swift binary on MLX and Metal, no Python, offline. Works with Claude Code, Codex and Ollama or OpenAI clients.. It is open-source on GitHub.

Is slotstream free to use?

slotstream is open-source under the MIT license, so it is free to use.

What category does slotstream belong to?

slotstream is listed under other in the Claudeers registry of Claude-compatible tools.

7 views
★ 402 stars
unclaimed
updated 8 days ago

// embed badge

slotstream on Claudeers
[![Claudeers](https://claudeers.com/api/badge/slotstream.svg)](https://claudeers.com/slotstream)

// retro hit counter

slotstream hit counter
[![Hits](https://claudeers.com/api/counter/slotstream.svg)](https://claudeers.com/slotstream)

// reviews

// guestbook

0/500

// related in Other

🔓

符合nature论文学术表达和科研绘图的Skill

// otherYuan1z0825/⟨Python⟩★ 45,880◷ Apache-2.0[ claude ]
🔓

Anti-AI-slop design skill for Claude Code, Cursor, and Codex.

// otherNutlope/⟨CSS⟩★ 29,187◷ MIT[ claude ]
🔓

Open source Ghostty-based macOS terminal with vertical tabs and notifications for AI coding agents. Built for multitasking, organization, and programmability.

// othermanaflow-ai/⟨Swift⟩★ 27,391◷ NOASSERTION[ claude ]
🔓

Huashu Design · HTML-native design skill for Claude Code · Claude Code 里 HTML 原生的设计 skill · 高保真原型 / 幻灯片 / 动画 + 20 设计哲学 + 5 维评审 + MP4 导出 · Agent-agnostic

// otheralchaincyf/⟨HTML⟩★ 24,472◷ MIT[ claude ]
→ see how slotstream connects across the ecosystem