Skip to content
Resources > Product Articles

Anthropic’s Fable 5.1 and OpenAI’s Astra: What the Frontier Does for Decision-Grade Intelligence

By Daniel Campos, Distinguished EngineerSeptember 14, 2026
A scatter plot showing AlphaSense search achieves much higher evidence quality scores (0.42-0.44) than Plain vector index (0.27-0.29), with model choice having little impact within each method.

Figure 1: Performance of the frontier intelligence with the frontier context. Fable 5.1 and GPT-6 Astra push the state-of-the-art ahead but are highly dependent on the quality of their harness.

OpenAI and Anthropic recently shipped new frontier intelligence within days of each other: GPT-6 Astra and Fable 5.1. Every time a new frontier model lands, we ask the same three questions: Does it make the analyst’s day better? Does it find insights faster? What does it cost to run?

Since the models were released, we’ve benchmarked approximately 30 setups across scale and size on a fresh set of 198 queries modeled on real analyst workflows. Like the ones used for our August benchmarking, these queries are challenging, multi-step finance questions. Each model is evaluated by a checklist of what good looks like, and we score each system by how much of the checklist is covered.

What We Found:

  • Fable 5.1 posts the best evidence score we have recorded and ties Opus 5, our production model, with no statistically significant gap.
  • Fable 5.1 reaches that quality with 31% fewer searches, 43% fewer steps, and 47% fewer tokens, which offsets its higher per-token price and keeps cost per query close to Opus 5.
  • GPT-6 Astra lands ahead of prior OpenAI models. It asks more numerous and more varied questions than any OpenAI model before it and is the first to get better as the toolset grows.
  • The harness, not the model, is the biggest differentiator. Even with frontier models, plain vector search returns documents that average 1.4 years old vs. 26 days with the full AlphaSense harness, while using 1.3 times fewer tokens for 1.9 times lower cost per question.

Fable 5.1 Has Top Evidence Score

Figure 2: The frontier compared to Fable 5.1

Our prior work found that the quality of discovered content is the primary driver of workflow success — we call this measure the evidence score. In this evaluation, Fable 5.1 posted the best evidence score we have recorded, though it isn’t significantly ahead of Opus 5, our current production system. Fable was better than Opus 5 on 71 questions, worse on 74, and 53 were even.

Figure 3: The frontier compared to GPT-6 Astra.

Fable Achieves Its Quality With Half the Work

Figure 4: Fable’s median usage per question, indexed to Opus 5.

While achieving roughly the same quality as Opus 5, Fable 5.1 reaches that level with 31% fewer searches, 43% fewer steps, 47% fewer tokens, and 39% less wall-clock time. That efficiency offsets Fable's higher per-token price, so the cost per query lands close to Opus 5 rather than above it. For work that runs thousands of times a day, matching the incumbent at a fraction of the effort matters more than edging past it at twice the cost.

Figure 5: GPT-6 Astra’s median usage per question, indexed to GPT-5.6 Sol

Similarly, Astra has huge gains over GPT-5.6 Sol. It issues 40% more queries while using nearly half the turns, and it’s 16% faster.

The Frontier Researches Differently, Not Just Better

Figure 6: Median tool usage and content discovery per question, every model on the same AlphaSense search tools

Models in the new frontier use tools differently compared to their predecessors and each other. GPT-6 Astra asks far more questions than its predecessors (26 vs. 17 for GPT-5.6 Sol and 9 for Terra), while Fable 5.1 pulls back to 26 from the 40 of the more tool-hungry Opus.

You might assume that the model is asking one question in different ways, but it isn't. Astra queries are the most varied on the board, so it approaches a question from more angles rather than circling it.

GPT-6 is also the first OpenAI model that lines up several searches at once instead of one at a time, averaging 2.6 tools per step. Fable 5.1 also batches, but it reasons out loud more and finishes in about a third fewer steps. For reference, Opus 5 searches harder than anything we have tested, at 40 searches per question. While each step uses a whopping 769 words per step, its turn efficiency leads it to be about cost-neutral with Opus-5.

A Bigger Toolbox Only The Frontier Can Handle

Figure 7: GPT-6 Astra is the first OpenAI model that sees score improvement with increased tool surface area.

To evaluate how well new models can use smarter and more diverse tools, we test each model twice. First with document and snippet search alone: the equivalent of an analyst who can only read. Then with 18 tools, including company financials, expert call transcripts, live events, and the knowledge graph: the equivalent of an analyst who can also pull any figure and sit in on an interview.

With search alone, Astra and the three GPT-5.6 models are roughly matched. If all you need is a good search, the model barely matters. But hand over the full toolbox, and they become more distinct. Astra pulls ahead, while Terra gets much worse as more surfaces give it more ways to make mistakes.

Each new model generation is better at turning a richer environment into a better answer, but that improvement is invisible without that environment.

Same Amount of Evidence, but Very Different Quality

Figure 8: GPT-5.6 Sol answering the same questions two ways. Source mix and document dates are read directly off the documents, with no grading model involved.

Even with the latest models, the quality of research is most significantly affected by what tools and content the model has access to. With plain vector search and a simple MCP tool, a frontier model searches about as much, but what it finds is usually stale sources and non-primary sources. The typical document the plain search returns is 1.4 years old, while with AlphaSense’s harness it’s 26 days. Similarity is not recency, and the same model without an intelligent harness will not deliver the same quality.

Better Tools Are Cheaper To Use

Figure 9: The same models, given different surfaces, see wildly different token use, and cost can vary by 53%.

Better retrieval is also cheaper to run. A frontier model with web search burns tokens on retrieving documents you can't use. But with AlphaSense, it uses 1.3 times fewer tokens for 1.9 times lower cost per question, because when the first search returns the right material, there is no second, third, or fourth to pay for.

What The Frontier Models Moved

For the last two years, the limit on agentic research was the model. This limit is gone. Both GPT-6 Astra and Fable 5.1 can plan ahead, issue work in parallel, and stay on task. Fable matches its predecessor's performance with half the tokens; Astra can search more diversely and competently than any prior model. But that isn't what drives quality.

The model is no longer the differentiator. Give frontier models the same search, and they land within a few points of each other. What decides the outcome is underneath: whether the documents are primary and current, and whether the tools return something worth having.

About the Author
  • Daniel Campos

    Daniel Campos, Distinguished Engineer

    Daniel Campos is a Distinguished Engineer at AlphaSense, where he researches how to improve language model inference, search, and information discovery for investment professionals and other knowledge workers. Before AlphaSense, Daniel founded Zipf AI and held research roles at Snowflake, Neeva, and Microsoft. He holds a PhD in Computer Science, an MS in Computational Linguistics, and a BS in Computer Science.

Explore more

Frontier AI Models Need Frontier Context: Why the Smartest Model Alone Won't Win

Today’s bottleneck on answer quality is no longer raw model intelligence — it’s context. Learn what we’re building next at AlphaSense.
model quality vs. cost with AlphaSense Search

The Library You Can't Crawl: Open Web Search vs. Licensed Data

We tested 241 questions through AlphaSense's agentic stack, comparing open web search to our proprietary content. See why combining both beats using either alone.
A chart titled "Cost vs. Citable Evidence" shows that the full AlphaSense retrieval stack offers the highest citable evidence at the lowest cost, outperforming leading web search providers and open web options.

The Hidden Cost of Building Your Own AI Research Stack

Reusing AI models without a data warehouse means paying the same hidden tax on every query. See how AlphaSense removes that cost.
building your own ai research stack

Transform intelligence
into advantage

Develop bold strategies, seize opportunities,
and lead with clarity and confidence.