Inject a fault.
A catalogue of ways your system can get worse, delivered through your own adapter. One exit code is the whole contract.
A suite that stays green proves nothing by staying green. Mutational breaks your system on purpose and shows which scenarios notice.
Runs against
The real binary, compiled to WebAssembly and otherwise unchanged, with a real recording beside it: six Meta-World manipulation tasks swept in MuJoCo, 542 simulator runs.
====================================================================
MUTATIONAL REPORT
====================================================================
runner replay: 542 recorded runs
discrimination 31.9% (23 of 72 injected defects detected)
scenarios 6 scored, 0 skipped, 10 inconclusive defects
undeliverable 26 defects could not be expressed in the scenarios aimed at,
and are excluded: not exposed, and not missed either.
PER SCENARIO
? reach-v3 0.0% 0/10 flake 0%
drawer-close-v3 15.4% 2/13 flake 0%
door-open-v3 16.7% 2/12 flake 0%
button-press-v3 45.5% 5/11 flake 0%
pick-place-v3 50.0% 7/14 flake 0%
push-v3 58.3% 7/12 flake 0%
Press "Run it here" to load the binary (3.7 MB) and produce this report
yourself, from the recording, in this tab.
A catalogue of ways your system can get worse, delivered through your own adapter. One exit code is the whole contract.
Each scenario's own pass/fail checks are the oracle. Verdicts are held against that scenario's own flake rate, so noise is never read as a catch.
Keep, delete, or cannot tell - each with the evidence behind it. And it says plainly when it cannot tell.
Hazards
A catch counts for the check that actually failed - not for what a fault was aimed at. See which hazards rest on a single scenario.
$ mutational gate --fingerprint v3
against arm-v1
heavy_mug_far_edge
no longer catches weak-grip x0.5
no longer catches weak-grip x0.75
exit 1 - the corpus lost coverageRelease after release, the gate fires on what a build stopped catching - not on a rate that moved. An unanswered question is never a pass.
Record a sweep once where the simulator lives. Replay it anywhere into the full report - in CI, on a laptop, or in this browser.
Worklist
What to do to the corpus next, for whoever - or whatever - is improving it. Field incidents raise demand; only a sweep is evidence.
POST /v1/sweeps curl https://collector.mutational.in/v1/sweeps --header 'Authorization: Bearer mutational_...' --json @sweep.json
Everything the CLI does is an API call or an MCP tool. Post sweeps from CI, ask the gate from a pipeline, drive it all from an agent.
This tab's worker removed
In the browser your files never leave the tab: the worker that reads your files deletes its own network first. On the collector, adapter output is never stored.
World models generate the universe. Simulators execute it. Mutational decides what it proves - and runs the loop that keeps a corpus worth running, from the cloud, while your simulator stays where it is.
Anyone may propose. Only a sweep accepts.
What your corpus lacks, most urgent first - computed in the cloud by the same binary as the command line, from the sweeps you post. Field incidents raise an item. Only a sweep closes one.
On your own OpenRouter key, in a format your project publishes: a closed vocabulary with bounds, never a path or a command. Every proposal is a claim, and is shown as one.
Or an autopilot does, inside a budget of campaigns, scenarios and dollars that you chose. A model cannot approve what a model proposed.
Next to the simulator, behind your VPN. It dials out, checks each scenario against the format in its own checkout, writes it to one file, and runs the adapter its own project file pins.
accept answers from that sweep: earns its place, duplicate, or catches nothing. About half of what a capable model proposes tests nothing and looks fine. Those are withdrawn.
On your World Labs key, only the scenes that earned their place - told the kinds of thing in the scene and nothing else. A background for a camera sweep, never counted as the scene.
Any agent that speaks MCP: over stdio beside the simulator, or over HTTP to the cloud with an agent token.
A model reads the worklist, writes a scenario, and asks a sweep whether it earned its place.
The model operating the real robot gets three tools and no more: to say what happened.
{
"mcpServers": {
"mutational": {
"command": "mutational",
"args": ["mcp", "--project", "mutational.project.json"]
}
}
}
tools orient · next · sweep_start · job_status · accept · gate
triage · asserts · catalogue · history · field_add · ...
suite.json
Every CI build posts its sweep to one collector. Two builds line up as one series, whatever machine ran them.
API tokens
An agent's token adds. It cannot delete, or replace a sweep with one that catches less.
Give the model that keeps your corpus a token that cannot rewrite the history its work is gated on.
DELETE /v1/sweeps/swp_2fb1... 200 the sweep, its rows, its recording at rest, per run: key · injected · outcome · reason
What the adapter printed - paths, serial numbers - is stripped before upload. Delete a sweep and its recording goes with it.
A migration broke two hundred scenarios. Know which to fix, which to delete, and why.
A discrimination evidence document per release: which hazards have a test that can fail.
Every new policy is a new system. See what the benchmark stopped being able to catch.
A loop a model can run by itself, with a referee it cannot argue with.
API, MCP and the whole loop included. Models run on keys you bring. Free while we are in beta.
Playground
$0
No account
Look around.
Free
$0 /mo
Sign in with your email · no card
Ship with evidence.
Enterprise
Talk to us
For regulated and air-gapped teams
Run it your way.
Can't find what you are looking for? Write to us.
Whether your simulation scenarios would fail if the system under test got worse. It injects faults from a catalogue - a blinded sensor, a lagging picture, a weak grip - and watches whether each scenario's own pass/fail checks object. A scenario that no fault can fail costs compute on every merge and is evidence of nothing.
One script that runs a scenario and honours an exit code: 0 if it passed, 1 if it failed, 3 if the fault means nothing in that scenario, anything else if the harness broke. The fault arrives as JSON in an environment variable. It has been run through reference adapters for CARLA, Isaac Lab and Meta-World on MuJoCo, written this way.
Sweeps run where your simulator is, and scenario source never leaves it: the collector refuses it by name, along with your adapter's command line, its environment and any path. What --post sends is a summary - scenario names, scores, which faults were caught - and, if you ask, the recording with everything your adapter printed stripped out.
A hosted worklist needs more than a summary, and gets it only when you say so: mutational project push sends the fault catalogue, what each scenario asserts and the format a model may write in, and says what it is sending every time it runs.
In the browser, nothing leaves the tab. The worker that reads your files deletes its own network access before it accepts one, and this site's account API refuses any request larger than four kilobytes - there is nowhere here to send a file.
1,000 sweeps posted a month, 10 GB at rest - recordings, full records, pictures and project documents together - 10 API tokens, 365 days of history; 20 projects with 5 workers each, 300 campaigns, 500 proposed scenarios and 100 renders a month. Browser use is not metered. A sweep refused for a limit is still on the machine that ran it, and can be posted later from its recording.
Sign in with your email and download it from your dashboard: one static binary for macOS or Linux, on Apple silicon, Intel or Arm, with its checksum. There is an install one-liner for CI.
Yes. mutational mcp serves everything over the Model Context Protocol beside your simulator, and the collector serves the same tools over HTTP for an agent that is somewhere else. The rule both enforce is the product: a model may propose anything, and only a sweep accepts it. Give an agent an agent token and it can add to your history, ask for proposals and queue campaigns - and cannot delete, weaken, approve its own proposals, or spend your render credits.
You do, directly, on keys you bring: OpenRouter for proposing scenarios, World Labs for renders, and optionally TypeSafe for filing field notes. A key is kept sealed - we can use it and cannot show it, to you either - and is sent only to the provider it belongs to. Every call made with one is listed in your dashboard with its host, its purpose and what kind of thing was sent. Set a spending cap at the provider: it is theirs to enforce. A set of five proposed scenarios has cost about four to eighteen cents; a draft render about eighteen.
Delete any sweep through the API and its recording goes with it. Delete your account from the dashboard and the whole workspace goes: recordings first, then every row.
Sign in with your email, download the CLI, and find out what your green suite actually proves.