Artificial intelligence has officially hit a wall. For years, we watched as large language models (LLMs) decimated standardized tests, moving from mediocre scores to 90th-percentile performances on the Bar Exam, GREs, and specialized medical boards. But the era of easy benchmarking is over. The introduction of Humanity’s Last Exam (HLE) has redefined what it means to test a machine, and as of early 2026, the results are humbling for Silicon Valley.

Humanity’s Last Exam isn't just another dataset. Created by a massive consortium involving the Center for AI Safety and Scale AI, this benchmark was specifically designed to be the final frontier of academic evaluation. It represents the point where pattern recognition ends and true, deep expert reasoning begins. While humans with Ph.D.-level expertise in niche fields can still navigate these questions with roughly 90% accuracy, even the most advanced frontier models are struggling to break the 50% barrier.

The Death of MMLU and the Need for Harder Problems

To understand why HLE matters, we have to look at the failure of previous benchmarks like MMLU (Massive Multitask Language Understanding). By 2024, MMLU had become effectively useless as a differentiator. Top-tier models were all scoring in the 80s and 90s, leading to a phenomenon known as "benchmark saturation." We weren't measuring intelligence anymore; we were measuring how well a model could recall its training data.

When a model hits 90% on a test, it often implies the test is no longer challenging the limits of its architecture. Researchers realized that to see the gap between "really good autocomplete" and "human-level reasoning," we needed a test that could not be solved by simple internet retrieval or basic logic. We needed something that required a synthesis of highly specialized knowledge, often involving obscure ancient languages, complex mathematical proofs, and multi-step visual reasoning.

How Humanity’s Last Exam Was Built to Stump Machines

The methodology behind HLE is what makes it uniquely brutal. This wasn't just a collection of hard questions; it was a curated filter designed to exploit the specific weaknesses of LLMs.

The creation pipeline functioned like an intellectual gauntlet. First, experts from over 100 disciplines—ranging from ornithology to biblical Hebrew—submitted thousands of original, closed-ended questions. Then came the "AI filter." Each question was presented to three frontier models. If any of them got it right, the question was discarded. The survivors were then sent to an even more powerful tier of models. Only if six consecutive top-tier models failed to solve the problem did it qualify for the final HLE dataset.

This ensures that HLE is composed entirely of questions that current AI architectures find inherently difficult. It targets the "reasoning gap"—the space where a model has the information but lacks the cognitive flexibility to apply it to a novel, highly specific constraint.

2026 Performance Review: Claude 4.6 vs. Gemini 3.1

In our latest internal testing, we ran the current leaders of the pack through the HLE public set. The results show progress, but certainly not a victory.

Last year, models like GPT-4o were barely scraping 3% accuracy on these problems. Today, the landscape looks different, but the ceiling remains firm:

  • Claude 4.6 (Sonnet/Opus): In our tests, Claude 4.6 showed the strongest capability in the humanities and ancient language sections of HLE. It achieved an accuracy of 42.1%. Its strength lies in its ability to follow extremely long, complex instructions without losing the thread of the logic. However, it still falls into "hallucination traps" when asked to identify microanatomical structures in avian biology—a common HLE subject.
  • Gemini 3.1 Pro: Gemini slightly edged out Claude in the mathematical and physics modules, reaching 44.5%. This is likely due to its superior integration with external tools and code execution environments. When a problem requires writing a script to solve a combinatorial puzzle, Gemini is less likely to make a syntax error, though it still often fails the underlying logical premise of the question.
  • The o1-Series Successors: The latest reasoning-focused models from OpenAI have crossed the 48% mark. While this is a massive leap from the 8% seen in earlier iterations, it still leaves a 40-point gap between the best machine and a human expert.

What’s fascinating is where these models fail. They don't fail because they don't "know" the facts. They fail because HLE questions often require what we call "cross-modal synthesis." For example, a question might present a diagram of a specific 18th-century mechanical device and ask how a change in one gear's teeth count affects the output torque based on a specific, non-standard physical constant provided in the text. This requires visual parsing, mathematical calculation, and contextual physics reasoning all at once. Even in 2026, the "transformer" architecture often stumbles when these three domains must be perfectly synchronized.

The Multi-Modal Killer: Why Diagrams are AI’s Kryptonite

About 14% of the questions in Humanity’s Last Exam are multimodal, meaning they include figures, graphs, or complex diagrams. This is where we see the most significant drop in AI performance.

While models have become excellent at describing what is in an image (e.g., "a photo of a cat"), they are still remarkably poor at spatial reasoning within that image. In HLE, a diagram isn't just an illustration; it's a core component of the logic. If a model misinterprets the angle of a line in a geometry problem or the connection between two nodes in a chemical structure by even a few pixels, the entire answer falls apart.

Our analysis suggests that current vision-language models (VLMs) still treat images as a secondary input rather than a primary logical space. Until AI can "think" in spatial terms as well as it does in linguistic terms, HLE’s multimodal section will remain a graveyard for frontier models.

Subjective Critique: Is HLE Too Niche?

Some critics argue that Humanity’s Last Exam is less a test of intelligence and more a test of "obscurity." Is being able to translate Palmyrene inscriptions really a measure of AGI?

In our view, the answer is yes, precisely because of the way the questions are framed. HLE doesn't just ask for a translation; it asks for a translation that accounts for specific historical context and linguistic drift that isn't readily available in common web-crawled datasets. This forces the model to move beyond its "stochastic parrot" tendencies.

If a model can solve a problem that only 500 people on Earth understand, and it does so by reasoning through the steps rather than just retrieving a pre-existing answer from its weights, that is a genuine signal of advanced intelligence. HLE is essentially forcing AI to act like a researcher, not a student.

The Reality of the "Human Baseline"

The most telling statistic in the HLE paper is the human baseline. Human experts—people who actually spent their lives studying these specific subfields—score around 90%. They don't get 100% because the questions are legitimately difficult, even for specialists.

However, the gap between the human 90% and the AI 45% represents the current "Reasoning Moat." This moat consists of:

  1. Logical Consistency: Humans can verify their own work. AI still struggles with "self-correction" in a zero-shot environment.
  2. Specialized Intuition: Experts can often sense when an answer is "wrong" based on first principles. AI often follows a flawed logical path to its bitter, confident end.
  3. Novelty Handling: HLE questions are designed to be original. They aren't in the training data. This strips away the advantage of massive scale and forces the model to rely on its internal architecture.

Looking Ahead: The Road to 90%

Will AI ever ace Humanity’s Last Exam? If we look at the trajectory from 2024 to 2026, the progress is undeniable. We went from single-digit accuracy to nearly 50% in two years. However, the next 40% will be significantly harder than the first 40%.

To bridge this gap, we likely need more than just "more compute" or "more data." We need a shift in how models handle multi-step reasoning—perhaps through more integrated chain-of-thought architectures or by moving beyond the limitations of pure next-token prediction.

For now, Humanity’s Last Exam serves as a vital reality check. It reminds us that while AI is incredibly capable of mimicking human-like output, it has not yet mastered the depth of human-level expertise. The exam is still ongoing, and for the first time in a long time, the machines are the ones sweating.

Final Observations from the Field

If you're a developer or a researcher, the takeaway from HLE is clear: don't trust your model's performance on public, common-knowledge benchmarks. The real test of your system's capabilities lies in its ability to handle the obscure, the visual, and the multi-layered.

Humanity's Last Exam has set the new gold standard for what "smart" looks like in the age of AI. It’s a brutal, uncompromising, and deeply necessary tool that ensures we don't mistake impressive mimicry for true understanding. As we move deeper into 2026, all eyes will be on the HLE leaderboard. The day a model hits 90% is the day we truly have to start asking what's left for us to teach them.