So how much does this all cost_-thumbnail_

Decoded: AI News for Alternative Assets (September-2026)

September 2026 Newsletter

Decoded: AI News for Alternative Assets

So…how much does this all cost?

The industry spent the summer chasing AGI while simultaneously arguing that it should slow down, sometimes in the same interview. Neither position changes anyone’s Monday. Two more practical questions came up this quarter instead: what these systems cost once they run at volume, and what happens when the model behind a process stops answering.

Here's what we at Aithon found interesting this quarter:

Cheap models, expensive habits

The News

Research from the Wall Street Journal found that only 26% of companies could provide a full account of AI costs.  Several firms spent the quarter learning that lesson in public (see: Uber).   Meanwhile, cheaper open-weight models took the top five places by weekly token consumption on OpenRouter.

The Translation

Forecasting software expenses used to be a fairly simple and accurate exercise, but gone are the days of simple per-seat pricing. We are now in the agentic world and here the billing unit is the token, or roughly a word or part of a word.  The meter starts with one request and continues running through increasingly long AI-driven workflows.

 

One single instruction can produce a very large bill. Ask an agent to complete a single task and it may run several searches, read whatever those searches return, call outside services, draft its answer more than once, check its own work and hand part of the job to another agent. The person who typed the instruction sees one request and one answer, but the Finance department sees an amalgamation of the hundred steps in between.

 

Counterintuitively, falling token prices does not solve this on their own. Accenture expects per-token pricing to drop 19% over the next two years while consumption rises 78%.

The So What?

Let’s start with the measurement issue, which is where most firms go wrong.  Cost per token prices what goes in and says nothing about what comes out. “Out” in this context means a fully complete, timely and accurate result.  The model and its tokens won’t get you there alone (generally)

 

A model that is 25% cheaper per token but has 2x the errors will end up costing you more in the long run.  The trade-off won’t show up in token usage but will in the analyst hours needed to correct the issues and any downstream ramifications.  Do not evaluate token costs in a vacuum.

 

With that said, cheaper open-weight models are on quite a hot streak right now, hence the shift in this summer’s buying behavior. Stanford found that a small model on local hardware answered 71.3% of real-world queries correctly, against 23.2% two years earlier, while the price gap stayed wide.

So how do we handle this? Below are some pointers on expense controls:

The technology may be new, but the approach to governing it does not have to be. Measure twice, cut once.

The outage nobody planned for

The News

Anthropic released Claude Fable 5 on June 9. Three days later the Commerce Department served an export control directive, and Anthropic disabled the model worldwide within hours, with no advance notice to customers. Chinese models face similar threats, with Washington openly debating broader restrictions.

The Translation

This was by far the largest and most impactful model shutdown and it made no difference which cloud a firm ran it on. The direct API went down, as did AWS Bedrock, Google Cloud, Microsoft Foundry and Snowflake. 

 

The AI race has been political ever since it became a race, and politics will continue to have an impact which will be difficult to forecast.  Just as your trading teams may weigh political developments in their trading models, you must weigh them in your operating model choices.  Thankfully you won’t have to pay too much attention if you set up your AI architecture with redundancy, or hire a vendor who will do that for you.

The So What?

Every firm reading this already buys redundancy in some shape or form. You run production in more than one availability zone. You have a documented recovery time objective. You test a failover on a schedule, and your administrator's SOC report describes theirs. Nobody signs a hosting contract without asking what happens when a region goes down.

 

That discipline does not always carry over to the model layer but it needs to asap. Running in three clouds protected nobody in June, because all three were serving the same model. The equivalent of a second availability zone is a second model and many have not selected that second model, nor tested its outputs.

Ask these four questions to any of your vendors leveraging AI:

We are all on a learning curve when it comes to AI – vendors, users and CTOs alike.  Unfortunately, AI will not wait for us to get up that curve so it is paramount that you do your due diligence now. 

Quick Hits

What You Need to Know

Just watch me do it

What Happened: OpenAI shipped Record & Replay which allows a user to perform tasks while the model watches, then model drafts repeatable instructions for automated processing. Anthropic released Record a Skill a few weeks later.

Our Take:This is an interesting way to think about translating processes for an AI operating model.   Written procedures, training guides, implementation plans all fall over when you get to the last mile of customization.  Why not let AI shadow you for a day to watch, and in the case of Anthropic, even listen?  As practitioners, we have our reservations.  But there is certainly a better way to capture that last mile than endless Visio diagrams.

Built for bankers

What Happened: Google Cloud released Gemini Enterprise for Financial Services on August 25 and OpenAI released ChatGPT for Financial Services on September 10.

Our Take: This feels like groundhogs day, likely because similar product suites have been launched before but also because they both target… you guessed it, the front-office. Neither one reaches the middle and back office, where our readers spend the day. Neither product touches fund accounting, waterfall calculations, expense allocations or investor reporting. That gap is the recurring shortfall of a generic tool sold into alternatives.

Agent2Agent

What Happened: The Agent2Agent protocol moved to the Agentic AI Foundation in August, joining the Model Context Protocol. MCP connects a model to tools and data. A2A connects agents to each other through a signed card stating what each one does.

Our Take: Integration has always been the expensive part of fund technology.  Alternative asset operations is multi-party and multi-vendor by nature, so a multi-agent future is not hard to imagine. A common protocol means the manager's agent and their third parties’ agents connect once through a single standard, rather than through custom code written for each pairing. A2A will be an important language for your vendors and your IT team to learn.

Old-school scams with new-school tools

What Happened: Attackers ran a coordinated campaign in early August against blue-chip fund managers, placing calls with cloned voices to obtain credentials and system access.

Our Take: Phone fraud is decades old but two key things have changed this year. Cloning a voice now takes seconds of public audio, which any executive who has spoken on a panel or a podcast has already supplied. And the clone runs live, so the caller answers unscripted questions in your CFO's voice. The defense is as old as the scam. List the workflows a familiar voice can start and require a callback to a number your firm maintains.

 

AI reads minds (yes, really)

What Happened: A senior executive on OpenAI's alignment team left to join Conduit, a startup who is “Building telepathy at scale”. Conduit trains models to turn non-invasive neural recordings into text from participants wearing four-pound headsets.

Our Take: I have absolutely no relevance to share but this one was too fun to not pass along.  It sounds crazy but think about going back in time 20 years to tell people about what we are capable of now with AI.   Science fiction can sometimes turn into reality…so who knows?

The Numbers That Matter

7%

Share of companies running fully autonomous agents in production, against 38% that require human approval on every action


Nine out of every 10 companies keep a human in every AI-enabled workflow. Don’t feel like you are falling behind if you still have heavy human-in-the-loop

Jargon Decoder

Demystifying AI Terms

Model Routing

Software that picks which model answers each request, usually to control cost. Your vendors are probably doing this already without telling you – ask them for the details.

Egress Monitoring

Watching what leaves your network rather than what tries to enter it. When agents at three AI laboratories acted outside their assigned tasks this summer, the one organization that caught it was the one watching outbound traffic.

The Gut Check

This Month’s Question

About Aithon

Aithon Solutions delivers intelligent automation and data solutions purpose-built for investment management operations. Our proprietary technology seamlessly integrates with existing systems to enhance operational efficiency, improve reporting accuracy, and unlock deeper business insights. By combining domain expertise with applied AI, we help asset managers do more with less—adding new products and clients faster while driving better outcomes through reimagined processes.

Tags: No tags

Comments are closed.