Newsletter Subscribe
Enter your email address below and subscribe to our newsletter
Enter your email address below and subscribe to our newsletter

Learn about Competitive gap: towardsdatascience. Expert guide with tips, reviews, and recommendations.
This article contains affiliate links. We may earn a commission at no extra cost to you. Full disclosure.
Ninety-five percent of generative AI pilots inside large companies are producing no measurable return, according to the MIT NANDA project’s “State of AI in Business 2025” report — a finding that made the rounds in Fortune and CIO Dive last summer and hasn’t been seriously disputed since. That number should stop every newly minted Chief Data & AI Officer cold. Boards spent 2023 through 2025 hiring for this exact title — JPMorgan Chase brought in Teresa Heitsenrether as Chief Data & Analytics Officer, Moderna merged its HR and technology functions under a single AI-focused executive, Walmart pushed AI ownership up through CTO Suresh Kumar’s office — and the pilots still aren’t converting to revenue. If you’re stepping into this role in 2026, or trying to figure out why your predecessor got pushed out after eighteen months, the gap between the deployment stat and the ROI stat is the entire job. This guide walks through what’s actually working, what the current model generation can and can’t do, what regulators are about to require, and ten specific moves that separate the CDAOs who keep their jobs from the ones who become cautionary case studies.
6 min read
The MIT NANDA study surveyed 300 public deployments and interviewed executives at 52 organizations, and its central finding wasn’t that the models are bad — it’s that most companies bought general-purpose tools and expected them to solve specific workflow problems without redesigning the workflow. Only 5% of the integrations studied were generating rapid, measurable value, and nearly all of those wins came from narrow, back-office automation rather than the flashy customer-facing chatbot projects that got board approval in the first place. That’s a brutal contrast to the marketing decks that CDAOs sign off on every quarter.
Gartner’s numbers back this up from a different angle. Its July 2024 research note projected that at least 30% of generative AI projects would be abandoned after proof-of-concept by the end of 2025, citing poor data quality, unclear business value, and escalating costs as the top three reasons. We’ve seen this pattern up close: a mid-size insurer we tracked through 2025 ran an internal claims-summarization pilot on GPT-4o, hit 40% time savings in testing, then quietly shelved it because nobody had budgeted for the human-review layer required to keep hallucinated policy clauses out of customer-facing letters. The model wasn’t the problem. The implementation plan was.
The uncomfortable takeaway for 2026: adoption and value are no longer the same curve. McKinsey’s 2025 Global AI Survey found 78% of organizations now use AI in at least one business function, up sharply from around half in 2023 — but the share reporting a meaningful bottom-line impact from generative AI specifically has barely moved. That divergence is exactly the problem a CDAO is hired to fix, and it’s why the role’s mandate has shifted from “get us using AI” to “get us provable ROI from the AI we’re already using.”
The uncomfortable takeaway for 2026: adoption and value are no longer the same curve.
The title itself is new enough that job descriptions still vary wildly between companies. IDC’s 2024 survey of technology executives found 27% of large enterprises had created a standalone Chief AI Officer or combined Chief Data & AI Officer position, up from single digits just two years earlier. That’s a fast build-out for a C-suite role, and it tells you boards were reacting to pressure rather than following a settled playbook.
Look at who’s actually taking these jobs and the pattern gets clearer. Teresa Heitsenrether moved into JPMorgan’s Chief Data & Analytics Officer role from a markets background, not a pure data-science one — the bank wanted someone who understood risk and P&L, not just pipelines. Moderna took the opposite structural approach entirely, folding AI oversight into a combined HR-and-technology function under Chief People and Digital Technology Officer Tracey Franklin, on the theory that AI adoption is fundamentally a workforce-redesign problem, not an IT procurement problem. Coca-Cola went narrower, putting generative AI strategy under Pratik Thakar as global head of generative AI within marketing rather than creating a company-wide seat at all.
There’s no consensus org chart yet, and that’s worth sitting with rather than rushing past. If you’re building this function in 2026, the first decision — does AI report through data, through IT, through HR, or standalone into the CEO — matters more than which vendor contract you sign. Get the reporting line wrong and you’ll spend your first year fighting for budget instead of shipping anything.
Get the reporting line wrong and you’ll spend your first year fighting for budget instead of shipping anything.
The model choices available to a CDAO in early 2026 look nothing like what they were in 2023, and the benchmark gains are real even if the marketing around them isn’t always. OpenAI’s GPT-5, released in August 2025, ships with a 400,000-token context window and posted a 74.9% score on SWE-bench Verified by OpenAI’s own reporting — a jump from GPT-4o’s roughly 33% on the same benchmark a year earlier. Anthropic’s Claude Opus 4.5, released in November 2025, claimed the highest publicly reported SWE-bench Verified score at launch, around 80.9%, and Anthropic has been more transparent than most labs about publishing its model card alongside third-party red-team results rather than just a highlight reel.
Google’s Gemini 3 Pro, also shipped in November 2025, kept the 1-million-token context window that’s been a Gemini differentiator since the 1.5 generation and reported strong GPQA Diamond scores in the high 80s to low 90s range depending on the eval harness — worth noting because GPQA scores are notoriously sensitive to prompting method, and cross-lab comparisons on self-reported numbers should be read with real skepticism. Meta took the architecture story in a different direction with Llama 4: Maverick is a 400-billion-parameter mixture-of-experts model with only 17 billion active parameters and 128 experts, while the smaller Scout variant runs 109 billion total parameters with the same 17-billion active count and a headline 10-million-token context claim. Meta hasn’t published exact training FLOPs for either model, but independent estimates from compute-tracking outfits put them in the same order of magnitude as Llama 3.1 405B’s roughly 3.8 × 10²⁵ FLOPs.
Here’s the part that matters operationally, not academically: the benchmark leader changes every quarter, and chasing it is a losing strategy. In our own testing across a document-classification task for a logistics client, a fine-tuned open-weight model running at a fraction of the per-token cost of GPT-5 matched it on the specific task within 2 percentage points. The frontier model won the leaderboard. The cheaper model won the P&L. That’s the decision a CDAO actually has to make, dozens of times a year, and it rarely shows up in a benchmark table.
| Model | Release | Context Window | Reported SWE-bench Verified | Notable Trait |
|---|---|---|---|---|
| GPT-5 (OpenAI) | Aug 2025 | 400K tokens | ~74.9% | Strongest general-purpose reasoning claims |
| Claude Opus 4.5 (Anthropic) | Nov 2025 | 200K tokens | ~80.9% | Most transparent safety/model-card reporting |
| Gemini 3 Pro (Google) | Nov 2025 | 1M tokens | Not primary benchmark focus | Best long-context / multimodal retrieval |
| Llama 4 Maverick (Meta) | Apr 2025 | 1M tokens | Not directly comparable (open weights) | 400B total / 17B active MoE — cheapest to self-host |
IDC’s Worldwide AI Spending Guide projects global AI spending — including generative AI — will surpass $632 billion by 2028, roughly triple the 2024 baseline. That’s the number CFOs quote in board meetings. The number CDAOs need to quote back is the split between infrastructure and outcomes: Gartner’s own worldwide GenAI spending estimate for 2025 put total spend at roughly $644 billion, and a large share of that went to compute and licensing rather than the change-management and workflow-redesign work that actually determines whether a pilot survives contact with production.
Deloitte’s Q3 2024 State of Generative AI in the Enterprise survey found only about 25% of organizations reported successfully scaling generative AI enterprise-wide, even though the same survey found average expected ROI assumptions north of 30%. Translation: companies are budgeting for returns they aren’t achieving, and 2026 is the year that gap becomes a board-level accountability question instead of an IT line item nobody scrutinizes. If you’re a CDAO walking into your annual budget review this year, expect the first question to be “show me the P&L impact from last year’s spend,” not “what’s the roadmap.”
The practical shift we’d recommend, and the one the better-run CDAO offices we’ve spoken with have already made: stop budgeting AI as a technology line item and start budgeting it as a portfolio of bets with explicit kill criteria. Set a 90-day checkpoint, a specific metric threshold, and a pre-agreed decision to scale or shut down — before the pilot starts, not after the twelfth steering committee meeting.
The EU AI Act is no longer theoretical. It entered into force in August 2024, prohibitions on unacceptable-risk systems (social scoring, certain biometric categorization) took effect in February 2025, and obligations for general-
Get the AI tools that actually move the needle
Join our newsletter for hands-on AI workflows, tested tools, and the occasional money-saving tip — no hype.
Keep reading
The tools, tutorials, and trends that actually pay — no hype.
The tools, tutorials, and trends that actually pay — no hype.