Why we're publishing this
The model-release news cycle is broken. A vendor publishes a model card with eval numbers — sometimes against a custom harness, sometimes with no harness specified at all. There's a flurry of attention, a few benchmark threads, then everyone moves on. Months later a paper shows the model was contaminated, or fails on a category nobody tested. By then it's in production, deployed on the strength of the card.
We think procurement teams, deployment engineers, and security reviewers deserve a better substrate — one with three properties:
- Sourced. When we state a spec, it comes from the publisher's own card or paper, and we link it. When we don't have a number, we say so rather than inventing one.
- Grounded in the registry. Every release post names a model that actually exists in our catalog, and states the tier the registry actually shows — Quarantine, Hardened, or Sealed — not a tier we wish it had.
- Verifiable. Claims link to the live artifact page and to /verify, so the trust chain doesn't bottom out at "we said so."
What you'll see here
Mostly release posts: when a model reaches Sealed or Hardened, a short data sheet — what it is, the sourced specs, what tier it holds in the registry, and how to deploy it. Occasionally a substrate post about how tiers, witnesses, and the transparency log work. We avoid marketing for things that aren't running, vendor-quote PR, hot takes, and SEO listicles. The catalog answers "what are the best models"; the blog is for the changes to it.
What's next
The first three release posts cover models you can open in the registry right now: Qwen3-235B-A22B (Sealed), Llama-3.3-70B on Jetson Thor (Hardened), and DeepSeek-V3 (Hardened). Each links back to its live artifact so you can verify the tier and provenance independently.
Welcome.
If you have to defend a model selection to an ATO board, need to know what runs on an edge box, or just care about open-weight AI being deployed responsibly — this is for you. See you on the next post.