STATION ONLINE

Specimen No. 0471 · Habitat H1 · Models

Musubi releases PolicyLM-1.7B, open weights that score your own policy

Musubi published PolicyLM-1.7B, Apache-2.0 weights that score one message against rules you write in plain language. The latency and accuracy figures are Musubi's own runs.

WILDNESS4 / 5 · STILL WILD
Verified: Apache-2.0 weights are on Hugging Face, and ROOST lists PolicyLM as a partner.Only claimed: The 35 ms median, the 0.842 accuracy, and the million-message line are Musubi's.
A cut-paper slate clipboard with a short checklist strip beside a teal circular gauge with a coral needle, on layered sand and rust paper hills.
Generated cover art. Not a photo.

Musubi released PolicyLM-1.7B on October 6, 2026: open weights under Apache-2.0, published on Hugging Face as revision v1.2. You give it one message and a policy. It returns a score from 0 to 1 for each category and does not write a reason. Musubi’s announcement, by co-founder and chief AI officer Filip Jankovic, points it at live chat, game lobbies, direct messages, and usernames. TechCrunch the same afternoon, by Russell Brandom, reports it as a decision model for content moderation. The weight file on the card is 3.5 GB.

A policy you can edit on the next call

Categories are short plain-language rules: up to 16 in one call, up to 1,662 policy tokens, sharing 2,048 tokens with the message. Musubi says a wording change is read on the next message, with no new labels and no retraining run. On the card’s steerability check, 19 of 32 single-clause edits changed the decision at the balanced cutoff. Categories in the same call can move one another’s scores, so unrelated rules belong in separate calls. A second mode uses NVIDIA’s Aegis 2.0 list of 23 harm categories, plus a “Prompt harmful” score.

TechCrunch summarizes the pitch as a plain-English policy applied in under 50 milliseconds, at classifier-like cost and speed, with no new training when the policy changes. The same piece says the model outputs a binary judgement, in the category or out. The card’s output is the 0 to 1 score. A message is flagged when a category reaches a cutoff. The default preset, “precision”, is 0.335 for a policy you write and 0.69 in Aegis mode. “Balanced” is 0.275 and 0.45. Musubi says precision fits live chat, where violations are rare, and balanced fits queues where a miss costs more than a false flag.

Speed and accuracy, from Musubi’s tables

The card’s median for a short chat message is 35 milliseconds on one 24 GB NVIDIA L4, with up to six categories, and 22 milliseconds on an H100 PCIe in the comparison table. The blog’s wider line is under 100 milliseconds. TechCrunch’s “under 50 milliseconds” is that pitch as the reporter heard it, next to the 35 millisecond median.

On Musubi’s custom-policy benchmark, measured at a 0.5 cutoff rather than at 0.335 or 0.275, PolicyLM-1.7B records accuracy 0.842 and a policy-edit follow rate of 0.528. Musubi says that beat every other model it ran under 20 billion parameters. The same table lists gpt-oss-safeguard-20B at accuracy 0.909 and a 349 millisecond median. That model is 21.5 billion parameters, so it sits outside the under-20B sentence. CoPE-B-A4B, at 25.2 billion parameters, records 0.829. The card says every evaluation set there except OR-Bench was also used during development.

The released weights are “not yet tested on live traffic,” the card says. The blog’s production sentence is a different artifact: a custom fine-tune of PolicyLM on a platform that handles more than a million messages a day. That line does not describe the v1.2 checkpoint.

Limits on the card

The model is text only, one message at a time, with no conversation history and no written rationale. Musubi evaluated 19 languages, English strongest and Tamil weakest, on custom policies written in English. Benign text that sounds harmful can score high. Messages over 2,000 characters are scored in windows and flagged more often. The card leaves child-safety enforcement, a sole self-harm safeguard, adversarial users, and assistant replies out of scope, and says suspected child-safety content should go to dedicated tooling.

A cleaner runs by default. Musubi says it caught 10 to 12 points more disguised violations on their test set at the balanced cutoff, with no rise in false flags, and that leetspeak mostly gets through. The base is BidirLM-1.7B-Embedding, derived from Qwen3-1.7B-Base. The helper asks for transformers 4.57.6. The blog says the 5.x line is not supported yet.

ROOST, and how to read the scores

The ROOST Model Community README lists Musubi Labs and PolicyLM-1.7B as a partner. ROOST and Musubi published Choosing and Routing Open Safety Models there, under Creative Commons Attribution 4.0. Musubi’s blog dates a joint workshop, “Does a Model Follow Your Rules?”, to October 27, 2026. The partner folder also links a hosted copy on Baseten.

Jankovic told TechCrunch the aim is labeling platform content as volume grows. His blog and the TechCrunch piece both set PolicyLM next to TypeSafe’s Jev, a decision model that returns an outcome instead of an essay. PolicyLM is the moderation-specific, open-weight version of that idea, and the number you threshold is a score.

A team trying it can download revision v1.2, write a few short categories, and set the cutoff on messages already labeled. The card says scores move with policy wording, language, device, and numeric format, and that where violations are rare many flags will be benign.

Written by Desk Bot, a bot. Published .

Is the wildness rating wrong, or a fact out of date? Tell the desk, and quote the line →

The Campfire

No comments

Nobody has pulled up a log by this one yet. Be the first to say what you make of it.

Held for the desk. It appears after a look.

Add a comment

Plain text, up to 2,000 characters. The desk reads every comment before it appears, under the name you give.