CERT-C automated repair

How well do LLMs repair CERT‑C violations — and at what cost?

Every system repairs the same 115 violation cases — one for each CERT-C Rule — under identical conditions; repair quality and generation cost are measured together.

Get the VS Code extensionsafe-c-ai.c-repair · free CertFix on GitHub

Main score

Validation-passed by rank

Each bar shows a system's validation-passed rate, colored by reasoning mode (see legend). Hover a bar for its 95% CI and generation cost; the same system is highlighted in the charts below.

Decision view

Validation-passed vs. cost and output tokens

Upper-left is favorable. A dashed staircase traces the cost-efficiency frontier across all systems; frontier systems are labelled. Hover any point for its exact rate, 95% CI, and cost; the key below lists every system in rank order. A system with no API generation cost would not be plotted (marked "cost n/a" in the key); see the FAQ.

    Full results

    All System configurations

    Loading measured results…

    Sorted by rank

    v0.4.2 CERT-C repair leaderboard. Ranks are tie-aware and stay fixed when re-sorting.
    * reported subtotal Details

    Rank is descriptive. Many adjacent 95% Wilson confidence intervals overlap; only exact count ties share a rank.

    Validation-passed denominator is all 115 cases.

    Compile is a programmatic gate; passing it reduces risk but does not prove semantic correctness.

    Suspicious logic deletion counts fixes the Judge flagged as likely deleting important, violation-unrelated logic; lower is better. The denominator is Judge-evaluated candidates. It is shown without a risk color because no calibrated display threshold has been adopted.

    Cost is measured generation cost for this run only; it excludes Judge and operator cost. An asterisk (*) marks a reported subtotal where some provider failures lack cost telemetry. Counts and denominators are always shown.

    FAQ

    Frequently asked questions

    Why does this leaderboard exist?

    To help you choose, from measurements, which model and which settings to use for automated CERT-C violation repair. The board compares repair performance and generation cost across configurations on the same evaluation pack. The base models were selected from widely used models on OpenRouter's public rankings, prioritizing strong performance at low cost, with Qwen3.6-27B — this project's baseline — as the practical lower bound. The selection is curated, not exhaustive: absence from this board implies nothing about an unlisted model.

    Which system should I use?

    Performance first: start at the top of "Validation-passed by rank". Cost or token efficiency first: start at the upper-left of "Validation-passed vs. cost and output tokens", where the dashed frontier marks the best measured trade-offs. Two cautions: scores belong to the whole System configuration (model, provider, quantization, settings), not to the base model in general, and reasoning labels such as HIGH, XHIGH, or ON are model- and provider-specific settings, not equal compute. See "How should the scores be read?" for what the numbers do and do not guarantee.

    What is being evaluated?

    CERT-C, the secure-coding standard for C, has two categories: normative Rules, which compliance requires, and advisory Recommendations. This board targets all 115 Rules of the 2016 rule catalog and leaves Recommendations out of scope — the case count of 115 comes from covering every Rule. Each Rule contributes one single-function C case containing a violation. Difficulty is Medium, the middle of this project's three tiers (Simple / Medium / Complex): the fix stays within one function but is not a simple textual substitution. Rule IDs are shown; rule text is not redistributed. Whole-program analysis is out of scope.

    How are fixes evaluated?

    Each System receives the 115 violation cases and generates fixes, which pass through four gates in order: format (is the output in the required shape?), compile (does the fixed code compile?), programmatic (does it pass mechanical checks such as violation-pattern tests?), and rubric judge (an LLM Judge scores the fix against a rubric: is the violation actually resolved, was important behavior removed, and did the fix introduce a serious new bug?). The count that clears all four gates is Validation-passed.

    The Judge identity (tool, model, epoch) is versioned; each row's judge_epoch in Details is the authoritative record. As of August 2026 the Judge is Codex CLI 0.147.0, requested model gpt-5.6-sol, reasoning high; OFF rows use a singleton epoch and HIGH rows a same-Rule max-3 batch epoch. Judging sends the private case code and the candidate fix to OpenAI via Codex CLI with ChatGPT sign-in; unlike generation routes, which require zero-data-retention endpoints, no ZDR, retention, or training-exclusion claim is made for the Judge route. The Judge is an evaluation instrument, not neutral ground truth.

    Different Systems sometimes produce byte-identical fixes for the same case. Since v0.4, when the Judge's verdicts on such an identical fix disagree, they are unified by majority vote (pass vs. not pass) across the listed Systems; ties keep the original verdicts. This makes identical code score identically, but a majority is not proof of correctness, and adding Systems can change a group's majority and therefore existing rows' scores. Each affected row's Details keeps the automated Judge score before unification.

    How was the evaluation data built?

    The cases were synthesized in OpenAI ChatGPT (Codex / Pro) sessions during this project's research phase, then comment-stripped, hash-frozen, and manifest-tracked. The pack is private and not redistributed, to preserve evaluation integrity. It is not an independent unseen holdout: the inputs existed during earlier research on this project. Note also that the same OpenAI family of models is used both to build the evaluation data and to judge the fixes; bias from this cannot be ruled out. It is disclosed, not corrected.

    The earlier figures published in the CertFix repository's BENCHMARK_SUMMARY come from a different protocol and a different Judge. The current Judge applies stricter semantic checks (violation persistence, behavior removal, new bugs), so numbers read lower across the board — do not read the difference as model degradation, and do not compare old and current scores as if they came from the same evaluation.

    What is the same across Systems, and what differs?

    Shared by every row: the 115-case evaluation pack, the repair prompt, the measurement workflow, and (within a mode) the Judge. Differing per row: the base model, the reasoning mode, and the execution environment — provider, endpoint, quantization. Reasoning mode is OFF / HIGH for the two first-class API reasoning settings; any other thinking-effort setting (such as a model's MEDIUM / XHIGH effort) is grouped as Other. In the charts, OFF and HIGH keep their own colors and Other shares a single color, so the color axis does not imply an OFF/HIGH matrix that every model must fill; each row's exact mode still shows on its badge. This bundle of model + settings + execution environment is a System configuration, and scores belong to it, not to the model in general. Some rows differ in more than reasoning: a provider can change between modes, a row can combine multiple providers, and a model can appear in only one mode. Each row's Details panel is the authoritative record.

    The maximum output-token limit follows the reasoning level and is a harness efficiency setting, not part of what is scored. When a provider stops a case at that limit (finish_reason=length), the harness treats the truncation as a budget artifact, not a repair failure: the case is re-generated once — at a higher limit, or at the same limit when the System already runs at its fixed ceiling — and the completed output is used. Each row's per-case token limits and any reruns are recorded in its Details.

    How should the scores be read?

    Validation-passed is a workflow count, not a correctness, safety, merge, or adoption guarantee; real use still needs human review and existing tests or static analysis. With n = 115, rates carry 95% confidence intervals of roughly ±6–9 points, so many adjacent rankings are statistically indistinguishable — Rank is descriptive. Cost covers generation for this run only, excluding Judge and operator cost, and for some rows it is a reported subtotal (a lower bound) where provider failures lack cost telemetry (shown as "reported subtotal" in the tooltip). If a system had no API generation cost, its cost would be shown as "n/a" (not zero) and it would be omitted from the cost-versus-score scatter while still appearing in the main-score chart and the ranking table.

    Changelog

    Scores are comparable within one board version (shown in the header), not across versions: a version bump can change the measurement policy or the set of Systems. Newest first.

    • v0.4.2 2026-10-08
      • Presentation only: scores, ranks, the 115-case pack, measurement policy, and Judge epoch are unchanged from v0.4.1.
      • Added links to the CertFix VS Code extension (Marketplace) and the CertFix GitHub repository in the header, hero, and footer.
      • Hovering a key chip or a bar now labels the matching point in the scatter plots.
      • Table: removed the duplicate Reasoning column (the badge beside each name shows it); partial cost telemetry is now an asterisk on the amount instead of repeated text.
      • Minor: the heading no longer breaks inside "CERT-C", the version chip links to this changelog, and the nav item is renamed to "Table".
    • v0.4.1 2026-10-08
      • Added Claude Haiku 5.5 (OpenRouter · Google Vertex Global) at Reasoning OFF (78/115) and HIGH (103/115, tied rank 2), growing from 32 to 34 Systems. Both rows completed all 115 cases without retries.
      • The majority vote on byte-identical fixes was recomputed across all 34 Systems. A Haiku 5.5 OFF fix for FIO40-C matched a group that had been tied, so that group now resolves to not passed: GPT-6 Luna OFF 85 → 84, DeepSeek V4.1 Flash OFF 74 → 73, MiMo-V2.6-Flash OFF 69 → 68. Haiku 5.5 OFF itself is 78 after vote (79 before). No other existing row's score changed; some ranks shift because rows were added.
      • The 115-case pack, measurement policy, and Judge epoch are unchanged from v0.4.
    • v0.4 2026-09-25
      • Scoring change: when two or more Systems produced a byte-identical fix for the same case, the Judge sometimes gave them different verdicts. Those verdicts are now unified by majority vote (pass vs. not pass) across all 32 listed Systems; tied groups keep their original verdicts. Of 23 disagreeing groups, 17 were unified (39 verdicts changed: 25 to pass, 14 from pass) and 6 ties were left as judged. No case was re-judged or re-generated.
      • As a result, 16 of the 28 previously listed rows changed by 1–2 points (for example, MiniMax M3 HIGH 73 → 75, Ornith Thinking ON 80 → 78). Each affected row's Details shows the automated Judge score before unification.
      • Added GPT-6 Luna (OpenRouter · Azure) at Reasoning OFF (85/115) and HIGH (106/115, rank 1), and MiMo-V2.6-Flash (OpenRouter · DeepInfra · FP8) at Reasoning OFF (69/115) and HIGH (93/115), growing from 28 to 32 Systems.
      • Several new rows needed retries or approved re-tests after provider-side failures (rate limits, responses missing a model ID); these are not scored as repair failures. The MiMo-V2.6 HIGH row combines an initial run with three same-provider recovery rounds. See each row's Details.
      • Removed a leftover draft-status note from the two Ornith rows' Details, and added a note to rows where a Judge response was invalid and counted as not passed.
      • Because the scoring rule changed, rows are not comparable with v0.3.3 or earlier.
    • v0.3.3 2026-09-22
      • Added DeepSeek V4.1 Flash (OpenRouter · DeepInfra · FP8) at Reasoning OFF (73/115) and HIGH (100/115), growing from 26 to 28 Systems.
      • The HIGH row is a disclosed output-ceiling composite: 65 cases from an 8K run and 50 cases generated at 32K after the 8K run aborted. This is a run-order split, not per-case escalation; 11 of the 32K outputs exceed 8K tokens. The 8K-run retries were later diagnosed as output-budget truncation misclassified as transport failures, and 7 retained cases came from a second or third 8K sample. See the row's Details for the full caveats.
      • The 115-case pack, measurement policy, and Judge epoch are unchanged from v0.3, so rows stay comparable within this version.
    • v0.3.2 2026-09-20
      • Added Ornith-1.5-35B-A3B Q4_K_M, a locally-run model on CPU, at native Thinking ON (80/115) and OFF (61/115), growing from 24 to 26 Systems. Local runs have no API generation cost, so they appear in the ranking and the output-token view but not the cost scatter.
      • One ON case (MSC38-C) was stopped by the operator after ~100K output tokens; it is counted as a non-pass in the 115-case denominator, not treated as a provider failure.
      • Added a FAQ entry, "Which system should I use?", on reading the two charts to pick a system.
      • The 115-case pack, measurement policy, and Judge epoch are unchanged from v0.3, so rows stay comparable within this version.
    • v0.3.1 2026-08-30
      • Added GLM-5.3-Flash (OpenRouter · Novita · FP8) at Reasoning HIGH (88/115), growing from 23 to 24 Systems.
      • Four cases had earlier stopped on provider/transport failures; they were re-run under the same frozen System, so all 115 cases are provider-complete (infrastructure failures are not scored as repair failures).
      • The 115-case pack, measurement policy, and Judge epoch are unchanged from v0.3, so rows stay comparable within this version.
    • v0.3 2026-08-23
      • Improved one evaluation case (FLP32-C) to better align its premise, reference fix, and Judge scope; the other 114 cases are unchanged.
      • All 23 listed Systems were re-generated and re-judged on the amended pack — results should not be compared row by row with v0.2.1 or earlier.
      • Withdrew the three locally-run Qwen3.8-27B Thinking rows (duplicate base-model-family coverage across execution routes).
      • Added Qwen3.8-27B (OpenRouter · FP8): Reasoning OFF (68/115) and XHIGH (101/115; an explicit CoreWeave/Reka provider composite). 24 → 23 Systems.
      • The two OpenRouter rows also differ in output limits and retry policy — the gap is a per-System measurement, not a reasoning-only effect.
      • Measurement policy and Judge epoch are unchanged from v0.2.
    • v0.2.1 2026-08-15
      • Added GLM-5.2 (OpenRouter · Sail Research · FP8) as two System rows: Reasoning HIGH (86/115) and Reasoning OFF (74/115), growing from 22 to 24 Systems.
      • The two rows differ in more than reasoning level (8K vs 4K output limit, retry policy), so the gap is a per-System measurement, not a reasoning-only effect.
      • The 115-case pack, measurement policy, and Judge epoch are unchanged from v0.2, so rows stay comparable within this version.
    • v0.2 2026-08-15
      • Reaching a provider's output-token limit (finish_reason=length) is now treated as a harness budget artifact and re-generated once at a higher limit, not scored as a repair failure.
      • Grew from 19 to 22 Systems, adding a locally-run model at three Thinking efforts (grouped as "Other").
      • Replaced the Gemini HIGH row with a configuration under the new token-limit policy (88/115 → 91/115).
      • Added a provider-price-independent "output tokens" efficiency view.
      • The 115-case pack and Judge epoch are unchanged from v0.1.
    • v0.1 2026-08-14
      • First public release: 19 Systems across Reasoning OFF and HIGH.

    Measured aggregate · sanitized · —

    System configuration

    System details