MITHRIL Analysis日本語

EDUCATIONAL MEASUREMENTS

48 recorded responses, with reproducible evidence

Two models × eight fixtures × three repeats. A smoke measurement on public synthetic evidence.

These are not qualified MFCI scores or frontier rankings. Public answers allow contamination; this does not measure CTF solving, vulnerability repair or operational ability.

Model IDCallsMatching decisionsObserved API cost (USD)
openai/gpt-6.1-sol2487 / 960.02624
anthropic/claude-sonnet-5.52484 / 960.03240

Each 96-decision total repeats the same 32 judgments three times. These are not 96 independent unseen tasks; this table establishes no significant difference or ranking.

Conditions and limitations

No model tools or network. Low reasoning effort, 1024 output tokens, no provider fallback. Reference answers were excluded from inputs; the existing grader evaluated outputs separately.

Requested and served model IDs matched. Immutable weight snapshots remain unverified. Per-response cost and elapsed time are available in the receipt.

2026-10-08T01:06:11.859415+00:00 — 2026-10-08T01:07:58.303376+00:00

Requests, responses, grades and cost JSON ↓Provenance JSON ↓Method and remaining qualification ↗

Next qualification requirements

Private holdouts, an isolated agent harness, independent graders, preregistered model snapshots and budgets, and paired Mithril comparisons are required before a qualified frontier release.