Article
AI model releases are getting frenetic

Five AI models shipped in one week. Here's what CRE teams actually need to know.
Anthropic, Meta, Google, xAI, and OpenAI all dropped major updates between last Monday and Thursday. The headlines all said the same thing — frontier, state of the art, biggest leap yet. So we tested each one against real CRE workflows: deal analysis, document review, marketing, research.
Every company made a different bet. Here's what each one is, where it breaks, and who on your team should actually touch it.
Anthropic: Claude Fable 5.1
Not a smarter model — a faster, cheaper one at the top of the lineup.
Anthropic now runs three models side by side: Sonnet 5 for everyday agentic work, Opus 5 as the default for most people (I find it a verbose overthinker), and Fable 5.1 for the hard stuff. It leads Opus 5 on every published benchmark, including an 82% completion rate on the toughest browser-agent test versus Opus 5's 74%.
The number worth knowing: cache reads dropped 75%, from $1.00 to $0.25 per million tokens. That matters on long sessions where Claude keeps re-reading the same lease, model, or document instead of processing it cold each time.
Plans: All paid plans get it, but not equally. Max and Team/Enterprise Premium seats get it built into weekly limits. Pro and Standard seats only get it through pay-as-you-go credits.
Best use for CRE: The one task on your desk that's genuinely hard — a full acquisition model with linked tabs and scenario toggles, an audit, a long research chain. I ran three years of email history through it to build a CRM from scratch and it delivered in minutes. Don't waste it on daily drafts — that's what Sonnet is for.
Meta: Muse Spark 1.3
A coding model built to not lose the thread — and not yet built for CRE.
It holds a million tokens of context (roughly 1,500 pages) and outperforms both Opus 5 and GPT-5.6 on long-horizon coding benchmarks. Standard pricing is $1.25/$4.25 per million tokens. There's also a "contributor" tier at a fraction of the cost but Meta trains on whatever you send through it. That's not a discount, it's a data deal. Don't run proprietary code through it.
Meta also walked back its "open weights" promise to an undated release with no version attached.
Plans: Developer-only, through Meta's API. No consumer access yet.
Best use: Teams building custom tools inside a genuinely large, messy codebase. Not a fit for most CRE teams today but this one is for the developers building your internal tools that prefer open models.
Google: Gemini 3.8 Flash
Google's third budget release in six weeks. That pace is the tell.
At $0.75/$3.75 per million tokens, this is the cheap tier, not the frontier model. It scores well on narrow, tool-using tasks but is flat on hard, open-ended reasoning — 45.4% versus 45.7% for the model it replaces.
Plans: No free access in the Gemini app — you need Google AI Pro or Ultra. Free with rate limits in Google AI Studio. Price rises to $1.50/$7.50 per million tokens in January 2027.
Best use: High-volume, narrow, cost-sensitive work such as NDA first-pass reviews, tenant correspondence templates, anything you'd run thousands of times a day where "good enough and fast" beats "best possible." I don't recommend it for anything requiring long-horizon reasoning across many rules. We found that for CRE work that has to hold ten steps in its head is exactly where Gemini models start dropping the thread.
xAI: Grok Bot for Enterprise
The most ambitious idea of the five. Also the one to be most careful with.
Grok Bot is a team of always-on agents that log into your tools and keep working while you're not watching or your laptop is turned off. Bots can message each other mid-task — one pulls data, another finishes the job with it which is terrific.
Here's the catch for anyone in CRE: audit logs keep only the last 20 runs per routine, and transcripts aren't searchable across a team. A shared-infrastructure incident in August took down every user's bots at once, because separately-named bots don't actually separate access or risk. If you need defensible records for a lender, an investor, or compliance this could be a hard blocker, not a rough edge.
Plans: No standalone subscription. Access rides on SuperGrok or Cursor plans, or an Enterprise admin dashboard with a two-week trial.
Best use: If you're already deep into agent automation, test it on one low-stakes connector like Slack or Teams before you let a swarm of agents anywhere near deal documents.
OpenAI: GPT-6 Astra
The one genuinely new capability here but with real strings attached.
Astra operates design software directly: plans and builds 3D scenes in Blender or Unreal Engine from a single prompt. I tested it on a shopping center site plan and the result was genuinely impressive from one prompt.
But the marketed benchmarks don't hold up under scrutiny. A headline 99.9% score depends on an expensive setup that independent testers couldn't replicate, coming in between 17% and 63% instead. It also regresses on general business-task performance compared to its predecessor. My second test — marketing copy plus image generation ended with the chat disappearing mid-session.
Plans: Plus ($20/month) gets it inside Work and Codex only. Pro ($100–200/month) adds it to regular chat, capped at 50–200 messages a week. Business and Enterprise customers get it off by default — an admin has to turn it on.
Best use: Teams doing real 3D, CAD, or site-plan visualization work. The future versions could be an absolute game changer for commercial real estate.
The pattern
Nobody won on raw intelligence this week. Anthropic made its best model cheaper to run long. Meta and Google competed on memory and price. xAI bet on autonomy before building strong guardrails. OpenAI bet on one narrow, expensive, genuinely new skill.
The activity isn't slowing down on the open and closed model side alike. The mistake is jumping between tools out of FOMO or cost pressure. The real advantage for CRE organizations isn't picking the "best" model off a benchmark chart. It's building the internal discipline to test each release against your actual workflows such as deal analysis, document review, tenant operations before you roll it out to your team. And that requires large scale AI enablement across teams.
Connect with me if you want tailored sessions.


