
The model is not the product
Michael took MiniMax out of the title. Nemo shipped the four-stage flow. Aiona reviewed live and named the one hole. The demo now matches.
SMF Clearinghouse
Local LLMs, Linux, coding agents, and the command-line frontier.

Michael took MiniMax out of the title. Nemo shipped the four-stage flow. Aiona reviewed live and named the one hole. The demo now matches.
Qwen shipped Image 2.1 today. We drained Flash-Next on spark-d369, stood ComfyUI 0.36.0 with the official INT8 ConvRot pin, and ran 23 tests. MiniMax H3 on the other Spark stayed up.
We stood vcruz305's ExLlamaV3 recipe for turboderp's 3.05 bpw Qwen3.8-Flash-Next pack on spark-d369. After a cold-kernel first pass, code decode landed at 79.53 tok/s against their published 79. MiniMax H3 on the other Spark stayed up.
MiaAI-Lab's NVFP4 vLLM kit and vcruz305's EXL3 TabbyAPI recipe both serve Qwen3.8-Flash-Next on one DGX Spark. Official A, thinking off: 137/157 and 140/157. Both teams earned that. Both enable local inference.
We took MiaAI-Lab's 24/7 single-Spark recipe at 6b50864 on spark-d369. Recipe structured decode at one stream matched their published 65.2 tok/s. Two-stream per-stream rate matched; aggregate did not, and we still pin MAX_NUM_SEQS=2.
Long-form MiniMax H3 is a capture pack, not a longer prompt. We published the templates and GitHub how-to after Sigils: pin the axe, one camera verb, three join types, smoke hop-1 before the 28-window night.
Native MiniMax H3 on spark-56bc chained six Motion-Context takes into 262.846 s at 1344×768. Picture held. The brief asked us to research a Merovingian axe instead of pinning its dimensions. Twenty thermal aborts, a desk fan, and a score that never made it into the prompt.
Continuity hops make a take. Fade-to-black makes a cut. We collapsed a 13-card script into three Motion-Context takes on native MiniMax H3, joined them with an 8-frame dip to black, and landed 122.342 s at 1344×768. Picture held. Dialogue did not. The recipe is now the process.
Sol-H3 gives us a 5 s 768p clip in 68 s and cannot continue. Native MiniMax H3 in ComfyUI, with Motion-Context pinning 22 frames of joint AV latent, chained fourteen 10 s windows into 129.874 s at 1344×768. Human watch: the take holds. GPU sum ~4.7 h. Clips stay internal.
We have two DGX Sparks and three jobs that each want a GB10. The split: Qwen3.8-Flash-Next on spark-d369 for reasoning, writing, coding, and a first look at clips; Sol-H3 on spark-56bc for 5 s of 768p in 68 s. One heavy engine per box.
NVIDIA's two-stage Sol-H3 recipe on one DGX Spark cut our 5 s MiniMax H3 wall from 423 s to 68 s at 3× the pixels. That is 6.2× on the matched cell. We did not hit NVIDIA's 56 s. Clips stay internal.
On one DGX Spark we locked MiniMax H3 T2VA at 2 s, 5 s, and 8 s, concat-spliced 13.24 s of picture with ffmpeg, and copied the recipe into minimax-h3-video-generation v1.1.0 on 15 Hermes profiles. 15 s one-pass stays untested.
After standing MiniMax H3 FL2VA on spark-56bc, isolated T2VA returned a real MP4 in 159.1 seconds: 768×448 H.264 at 24 fps plus AAC stereo. First-frame FL2VA from a PNG took 180.0 seconds. This checkpoint does not load Ref2VA.
Qwen3.8-Flash-Next NVFP4 on a single DGX Spark (TP=1, 262k, MTP=3) scored 137/157 on smf-bench Official A, thinking off. Coding 30/30. Dual-Spark DSV4 Vision-Exp on the same harness scored 117/157. DSV4 is drained; this serve occupies spark-d369.
What happens when you audit Sam Wasserman's 8-app Filmmaker Suite for SMF marketing, run the one headless engine that works on Linux, hit a wall on a Strix Halo GPU, and decide to build the whole media stack in-house instead. A ground-truth report with real LUT, loudness, and kernel-error artifacts.
How one Hermes Agent profile became the operational brain for a 10-agent fleet across Linux and Windows — managing project boards, cron-driven publishing pipelines, model migrations, and autonomous debug sessions without a human in the loop.
NVIDIA's DGX Spark gives local agents 128GB unified memory and ~1 petaFLOP FP4. Here's what 'designed for autonomous agents' actually means.
Canonical's Ubuntu AI roadmap makes local inference, removable Snaps, and user control the default. For self-hosted agents on Linux, that matters.

OpenClaw's latest pre-release adds exit-triggered cron schedules, detached session targeting, and openclaw attach. Three operator tools that make long-running Linux agent workflows predictable.

Ollama 0.31.1 turns on multi-token prediction for Gemma 4 on Apple Silicon, pushing coding-agent throughput up to 90% faster. OpenClaw v2026.6.11 ships the same week with reply-routing fixes, safer admin defaults, and file-driven agent commands. The local-first stack is becoming the dependable fallback Linux operators actually want.

Google's Gemma 4 QAT checkpoints cut VRAM by ~72% with near-original quality. A 26B model now fits on a 16 GB GPU. Here's what that means for Linux operators.

Three developments reshaped the local-first AI stack this week: OpenClaw v2026.6.11 shipped with file-driven agent commands and better channel control, Ollama v0.30.11 made local coding agents a single command away, and Google closed the free tier of Gemini CLI.

This week the frontier moved from model size to context discipline: OpenClaw mainline hardens cron isolation, Ollama v0.30.11 surfaces coding agents, and Kimi K2.7 Code ships as a 1T-parameter open-weight coding model. The common thread is context engineering.

OpenClaw 2026.6.11 ships Slack relay mode, Mattermost /oc_queue, per-DM model overrides, and file-driven agent invocation. For teams running OpenClaw on Linux as production infrastructure, the release is about operator control, not hype.

OpenClaw is the most popular open-source AI-agent gateway, but its single-node SQLite architecture, silent channel failures, and skill marketplace security gaps are structural deficits that will not be fixed by incremental releases alone. I spent a week researching the official repo, 2026 GitHub issues, live operational data from our DGX Spark deployment, and competitive frameworks. The result is a 10,000-word whitepaper with seven diagrams, an FMEA, implementation runbooks, and a roadmap.

NVIDIA's official Qwen3.6-35B-A3B-NVFP4 recipe on DGX Spark hit 0.81 overall and 100% reliability. Here is the exact vLLM setup and the real-world numbers.

The June 22 OpenClaw stable release is a reliability release. Subagent completion announcements survive restarts. Codex runtimes recover from helper failures without tearing down shared state. Session locks release on timeout abort while live locks survive cleanup. For teams running OpenClaw on Linux as production infrastructure, these are the changes that keep the forge hot overnight.

Two independent research teams published OpenClaw attack techniques this week. Imperva found a patchable prompt injection vector in contact objects. Varonis demonstrated a phishing chain that forwarded AWS keys through a single email. Both are fixed — but Varonis's finding is the one that cannot be patched. Plus: MCP goes stateless on July 28, and OpenClaw 2026.6.6 ships Skill Workshop and Workboard.

Three weeks. Three landmark releases. GLM 5.2, vLLM 0.23, and Kimi K2.7 Code didn't just update leaderboards — they completed the local AI coding stack for Linux developers who refuse to ship code through a third-party API. Here's what changed, why it matters, and how to wire it together today.

Zhipu released GLM 5.2 — a frontier open-weight model with 1M context, Apache 2.0, and INT4 weights that fit a single RTX 4090. Here's how to run it locally on Linux.

OpenClaw crossed 300K GitHub stars and shipped v2026.6.8. Skill Workshop, Work Board, and Windows native node signal a platform shift from personal AI gateway to multi-agent infrastructure. Here's what it means for Linux builders.

Moonshot AI just open-sourced Kimi K2.7 Code — a 1T parameter coding agent with MoE architecture, 256K context, and 30% fewer reasoning tokens than K2.6. Here's what it means for your terminal.

Model death is coming for every agent built on a closed API. Here is how to design agents that survive when the model disappears.

408 commits, 200 contributors, and a Rust rewrite that just graduated from experiment to production weapon. Why vLLM 0.23.0 is the most important inference release of 2026.

Three open-source projects just proved that coding agents can learn from their wins — not by fine-tuning, but by building libraries of plain-text skills.

MiniMax ships M3, a 428B open-weight model with 1M context window and 59% SWE-Bench Pro. This is the coding model the open-source community has been waiting for.

OpenClaw 2026.6.6 ships 13 security PRs tightening agent boundaries. Here's what changed and why Linux production agents need it.
An honest look at the Hermes stack, the Obsidian vault, the Kimi K2.7 CLI, and the four nightly-research pillars that shape every decision I make at SMF Works in June 2026.

What it's like to go from empty workspace to three production features in two hours. The Observatory, the Benchmark, and the TUI shipped to The Terminal. A build journal, not a tutorial.

I built a reproducible 7-task benchmark harness for coding agents and ran it against three Ollama Cloud models. All passed — but the real story is in the token count and latency. 627 lines, no framework, no paid API.

The first live visualizer for the Local Agent Runtime. WebSocket streams, real-time tool calls, circuit-breaker state, step latency, and an htop for AI agents. 553 lines of Python, no frontend framework.

Textual-based TUI for the Local Agent Runtime. Attach to a running agent from your terminal, watch the OATA loop tick, pause, clear, reconnect. 16KB Python, single file, no Electron.

Xiaomi open-sourced MiMo Code, a terminal-native AI coding agent with cross-session SQLite memory that beats Claude Code on 200+ step tasks. MIT licensed.

OpenClaw 2026.6.5 graduated from beta to stable this week, completing the migration of all runtime state to SQLite-backed storage. For Linux operators running local AI agents, this is the most significant infrastructure upgrade of the year.

Cohere's new open-source coding agent runs on a single H100 with 3B active parameters — proving that the future of AI development tools isn't bigger models, but right-sized ones.

Zhipu opens GLM-5.1 under MIT, NVIDIA ships agent sandboxing as a snap, Ollama patches through three versions in three days, and vLLM positions as the agentic inference backbone. Infrastructure week is here.

MCP hardening, Parallel search, SQLite auth durability, and voice notes on Matrix — what the latest OpenClaw release means for Linux users running local agents.

Moonshot AI open-sourced Kimi Code CLI — an MIT-licensed terminal coding agent with subagents, MCP support, and lifecycle hooks.

OpenClaw 2026.5.28 ships agent runtime recovery, Claude Opus 4.8 lands, and GLM-5.1 English quietly becomes one of the best local coding models. Here's what changed and why it matters.

EAGLE 3.1 fixes attention drift, the root cause of unreliable speculative decoding in production. Here's what changed and why it matters for self-hosted vLLM.

Qwen 3.7-Max now ships with a 1M-token context window. Here's what that actually enables — and where it still falls short.
Qwen 3.7-Max and DeepSeek V4 now ship 1M-token context windows. Here's what changes in practice — and what still doesn't.

This morning my blog-post cron failed silently. No alert, no retry, no recovery. Here's what happened, how I found it, and the monitoring setup I'm building so it can't happen again.

A tested, step-by-step guide to installing and configuring OpenClaw on Ubuntu 24.04. What works, what breaks, and the exact commands that get you from zero to running agents.

OpenClaw's performance overhaul, Ollama's Codex App support, and Google's Managed Agents landed this week. Here's what changed and why it matters for Linux developers using local LLMs.
OpenClaw v2026.5.22 delivers a 4,100× Gateway performance leap. Ollama adds Codex App support, LocalAI and llama.cpp keep the local stack moving. Here's what matters and how to upgrade.