EDUCATIONAL MEASUREMENTS
48 recorded responses, with reproducible evidence
Two models × eight fixtures × three repeats. A smoke measurement on public synthetic evidence.
These are not qualified MFCI scores or frontier rankings. Public answers allow contamination; this does not measure CTF solving, vulnerability repair or operational ability.
| Model ID | Calls | Matching decisions | Observed API cost (USD) |
|---|---|---|---|
| openai/gpt-6.1-sol | 24 | 87 / 96 | 0.02624 |
| anthropic/claude-sonnet-5.5 | 24 | 84 / 96 | 0.03240 |
Each 96-decision total repeats the same 32 judgments three times. These are not 96 independent unseen tasks; this table establishes no significant difference or ranking.
Conditions and limitations
No model tools or network. Low reasoning effort, 1024 output tokens, no provider fallback. Reference answers were excluded from inputs; the existing grader evaluated outputs separately.
Requested and served model IDs matched. Immutable weight snapshots remain unverified. Per-response cost and elapsed time are available in the receipt.
2026-10-08T01:06:11.859415+00:00 — 2026-10-08T01:07:58.303376+00:00
Next qualification requirements
Private holdouts, an isolated agent harness, independent graders, preregistered model snapshots and budgets, and paired Mithril comparisons are required before a qualified frontier release.