Skip to content
SUNDAY, AUGUST 2, 2026
AI & Machine LearningLegacy Report1 recorded source

Errorquake reveals heavy tailed LLM errors

Visual status: no verified article image is available. The reporting remains text-first.

Matched accuracy hides how LLMs fail, in brutal detail. The paper shows that at the same reported accuracy, open-weight models diverge dramatically in both the fault sketches and the consequence of those faults, not just in a single error count.

The researchers introduced Errorquake-10k, a 10,000-query benchmark that scores each response on a continuous 0-4 severity scale across eight domains and five difficulty tiers. They fit per-model severity distributions for 21 open-weight models and estimate a severity distribution index, b, with 95 percent bootstrap confidence intervals. The aim is simple in intent but hard in practice: measure not just how often a model goes wrong, but how badly it goes wrong. The headline claim is that severity carries information that a lone error rate cannot capture.

The study reports striking cross-model differences. Across 210 model pairs, 85 pairs have disjoint 95 percent confidence intervals on the severity index at matched accuracy. That means two models can look equally capable by a single score, yet produce very different distributions of harm when they err. The example cited pits deepseek-v3.2 against ministral-14b at epsilon = 0.586, with a Delta b of 0.47, illustrating how two apparently similar performers can diverge in the tail of failures. In other words, the same accuracy can hide a meaningful risk structure.

To back that claim with human judgment, the team reports a 519-item, three-rater validation study. It yields an ICC(2,k=3) of 0.85 for measurement reliability, a ranking correlation with human judgments (rho) of 0.89, and a human-to-model alignment measure (rho_s) of -0.86 relative to dense-model scaling. In short, the automated severity scores align well with human judgments, even as they reveal variance that accuracy alone misses. The results support a broader claim: severity distributions correlate with model size and error type, and those shifts are systematic rather than random.

A formal Non-Reducibility Theorem seals the point: the severity profile and the error rate are informationally non-redundant, with I(b; model | epsilon) = 1.56 bits. More than half of cross-model variance in the severity index remains unexplained by accuracy alone, specifically 64.5%. A severity mechanism taxonomy with kappa = 0.83 further clarifies how error types transform with severity: low-severity errors skew toward retrieval (71%), while high-severity errors skew toward fabrication (39%), and this composition shifts with model size (p < 0.0001). The takeaway is stark: better accuracy does not guarantee fewer or less dangerous failures, and models of different sizes can fail in qualitatively different ways.

For practitioners, the implications are concrete. The paper shows that severity-aware evaluation should accompany accuracy, and that the distribution of errors matters for risk management. Here are key takeaways researchers and product teams can act on:

  • Rethink success criteria. Don't rely on a single accuracy number; pair it with a severity distribution index and domain-difficulty breakdown to reveal how failures actually hurt users.
  • Benchmark with intent. Adopt a severity-aware benchmark like Errorquake-10k to compare models across multiple dimensions, not just aggregate error counts. The fact that 85 of 210 model pairs show non-overlapping severity intervals means apples-to-apples comparisons require richer metrics.
  • Align with human judgment. Confidence in automated severity scores increases when they correlate with human evaluation (ICC and rho figures reported), and when they reflect user-relevant failure modes such as fabrications.
  • Prepare for heavy tails. Design product gates, monitoring, and fail-safes around high-severity errors, which dominate risk even if they are rarer, and watch for how error composition shifts with model size.
  • Ultimately, the study argues for a practical shift: report severity alongside accuracy, because severity carries discriminative information about risk that plain error rates cannot capture.

    Sources & methodology
    1. ERRORQUAKE: Heavy-Tailed Error Severity Distributions in Open-Weight Large Language Models
      arXiv ML / Primary source / Published JUN 04, 2026 / Accessed JUN 05, 2026

    Newsletter

    The Robotics Briefing

    New signups are closed while external email delivery is being verified. No email address is collected here.

    Follow the live RSS feeds