Diary of an AI Architect

Diary of an AI Architect

GPT-6 Astra vs Claude Fable 5.1: what’s different and how I’d use them in production

Beyond the benchmarks: a practical guide to choosing the right model for each task, setting autonomy limits, and measuring the cost of getting it right.

Anurag Karuparti's avatar
Anurag Karuparti
Sep 19, 2026
∙ Paid

What a week in AI. Is artificial general intelligence (AGI) already here? And are we doing enough to keep increasingly capable systems safe?

This week in my diary, I’m bringing those questions back to the work: building AI systems that are useful, reliable, and stay within the boundaries we set.

As AI architects, we have an increasingly important role in shaping how these systems are designed, deployed, and used across society, while understanding and mitigating the risks they may pose to humanity.

On a personal note, I was promoted to Principal AI Apps Architect (Director) at Microsoft. It’s been an incredible journey, and I’m grateful for the people who have helped me grow.

I also earned Anthropic’s Claude Certified Architect - Professional certification. I took it to sharpen my architecture judgment when building multi-agent systems, not just learn another product.

What I valued most was the mix of technical principles and broader architectural decisions: what to automate, how to measure success, enterprise integration, and where human oversight belongs, GTM strategies and stakeholder enablement.

These are principles I can apply well beyond Claude, across the models and AI systems I work with, when enabling my customers.

In my next post, I’ll share how I cleared the certification and why I liked it so much.

For now: GPT-6 Astra or Claude Fable 5.1? Let’s look at execution, building, time, and cost to find the right fit for your business. The evidence comes from published evaluations, my first-hand research playing with both models.

I created a YouTube Short too on this. Check it out here!


The architect asked to pick one winner

Imagine Priya, a staff architect at a national retailer. Her finance team wants an agent to open a sales spreadsheet, reconcile the numbers, build a chart, update a presentation, and save the accompanying document.

Her engineering team wants something different: investigate a recurring inventory-service bug, trace dependencies, propose a fix, and verify it without breaking another service. Procurement asks Priya to standardize on one model for both.

The spreadsheet workflow fails if the agent updates the wrong file. The debugging workflow fails if a plausible patch hides the underlying defect. A quick answer is not the definition of done in either case.

Priya is not choosing the smartest colleague in the room. She is assigning responsibilities, tools, and acceptance criteria to two different jobs.


If you enjoy visuals like these and learn better visually, check out my Visual Library on Gumroad.


The framework: model loyalty versus workload fit

I use a four-decision routing test:

Reflex: Pick a model -> give it every task -> optimize the bill later
Fix:    Define the job -> test the model + tools -> verify the outcome
  • Execution: Can it complete the required actions across your tools?

  • Building: Can it understand the problem and produce a correct, maintainable artifact?

  • Time: Does it meet the deadline for an accepted result, not just the first answer?

  • Cost: What does that accepted result cost, including cache writes, retries, and review?

I use benchmark evaluations to decide which model to start with for each enterprise workload. I look at what each test actually measures, then connect that strength to the work our teams need done.

Astra

  • Astra feels like an operator, the get-things-done model. It outperformed Fable 5.1 on AutomationBench-AA, which tests completing multi-step business workflows across SaaS applications through APIs while following business rules.

  • That makes it my starting choice for updating CRM records, coordinating customer follow-ups, processing support escalations, employee onboarding workflows, and checking expenses against budget rules. [1][2]

  • Its stronger Terminal-Bench 4.0 results point toward command-line engineering and operations work: troubleshooting failed builds, fixing dependency issues, debugging CI/CD pipelines, diagnosing data or ML jobs, and running scripts and automated tests. These tasks require the model to actually operate tools and configure environments, not simply answer coding questions. [1][2]

  • For browser and desktop automation, I’d also prioritize Astra: retrieve supplier invoices, fill portal forms, update spreadsheets, perform frontend QA, or move information between enterprise applications.

  • OSWorld 2.0 tests operating that software, and OpenAI reports Astra ahead of GPT-5.6 Sol and Claude Opus 5 on its offline, partial-credit comparison. That supports my evaluation too, although the published setup does not establish a matched comparison against Fable 5.1. [3][4]

Fable

  • Fable feels more like an analytical builder. It outperformed Astra on SciCode, which tests scientific programming rather than software development broadly. I’d start with it for engineering simulations, numerical calculations, scientific data analysis, and research-oriented code where domain reasoning and code generation have to work together. [1][2]

  • It also leads Astra on AA-LCR v1.1, which tests reasoning across long documents. I’d apply that strength to comparing supplier contracts, analyzing requests for proposals (RFPs), reviewing financial reports, finding conflicting policies, mapping regulatory requirements, and checking requirements against technical documentation. [1][2]

  • Then I check the economics. Astra’s lower cost per task in the cited evaluation and Fable’s cheaper cached-input pricing create different opportunities.

  • Fable becomes particularly interesting when the same large context, such as policies, codebases, contracts, or documentation, is reused repeatedly. I’d compare reasoning effort, eligible cache reuse, latency, failure rates, and human-review costs.

  • The final choice depends on correct results, human effort saved, and total cost per completed workflow, not token price alone. [1][8][9][10]

  • The independent comparisons use Artificial Analysis’s September 15 snapshot at maximum reasoning effort, with Fable’s default fallback enabled, meaning another model may handle some requests.

  • These are workloads I would evaluate based on benchmark signals, not measured customer outcomes. [1][2]

  • Below, I unpack all four decisions, explain where smaller models fit, and show the handoff I would use to test a Fable-plans/Astra-executes workflow.

    You do not need a frontier model for every task.

    When I design an enterprise AI workload, I increasingly think about four decisions: what role the model needs to play, how much intelligence the task actually requires, how much autonomy I am willing to give it, and what it costs to produce a correct outcome.

    The goal is not to find the single best model. The goal is to use the right amount of intelligence at each step of the workflow.

Decision 1: Is the model planning or executing?

User's avatar

Continue reading this post for free, courtesy of Anurag Karuparti.

Or purchase a paid subscription.
© 2026 AgentChainAI LLC · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture