Latest
Cloud

The hidden decisions behind every AI request

The choice of model is only the beginning of an AI request. Take an insurance claim moving through an AI agent. It reads a photo of a damaged bumper, pulls up the policy, checks a repair estimate against regional pricing, drafts a summary and hands the file to a human adjuster. What this business counts […]

TechChallenger Staff9 October 2026 at 07:00 UTC9 min read
The hidden decisions behind every AI request

The choice of model is only the beginning of an AI request.

Take an insurance claim moving through an AI agent. It reads a photo of a damaged bumper, pulls up the policy, checks a repair estimate against regional pricing, drafts a summary and hands the file to a human adjuster.

What this business counts as one task is actually 40 or more calls to a model.

Each of those calls carries its own requirements. Reading a scanned form is simple work that a small model handles for a fraction of a cent. Judging whether the claim looks fraudulent needs a larger and slower model, and if the file contains medical records, those may not be allowed to leave the country.

For most of the past three years, enterprises have focused on which AI model to use. Now, the focus is turning to a broad question of placement: which model handles each call, where it runs, and whether there is capacity to take it.

That shift is showing up in budgets. Gartner reported in March 2026 that agentic models consume between 5 and 30 times more tokens per task than a standard chatbot exchange, even as the cost of running inference falls for providers. EY has put a figure on the same trend: a customer service interaction that cost about $0.04 to serve in 2023 costs roughly $1.20 in 2026 once it runs as an orchestrated agent with tools, reasoning steps and subagents.

The weight of AI compute is moving with it. Gartner forecasts that global spending on inference will pass spending on training this year, at $23.3 billion against $19 billion. McKinsey expects inference demand to grow at a 35% compound annual rate through 2030, by which point it becomes the dominant AI workload in data centers.

Three problems that used to belong to different teams now arrive together: which model does the work, where there is compute to run it, and who keeps it running once it is live.

“Companies spend months choosing a model, then find out the model was the easy part,” says Davy Wang, founder and CEO of GoodVision AI. “In production the same application makes thousands of small decisions every second on model, environment, and cost. These decisions determine whether AI pilots survive past the first year.”

GoodVision AI started in California in 2019, supplying and managing cloud capacity for enterprises in gaming, video and cross-border commerce. Since March it has also sold inference services, priced by usage, through a Smart Routing Engine that decides how each request is handled. A third business unit, called AI Factory, would put inference compute in compact sites closer to users. Wang worked at IBM, Amazon Web Services and Intel over 25 years before founding GoodVision AI.

Every request is a placement decision

Few companies run on one model any more. Research by S&P Global Market Intelligence, commissioned by Vultr, found the most advanced AI adopters running an average of 175 models in production and expecting to reach 200 within a year.

A single default model forces a compromise on every request. The largest model is expensive for work that a smaller one does well; Anthropic prices Claude Haiku 4.5 at a third of Claude Sonnet 4.5 and reports similar coding performance on some evaluations. Sending everything to the small model instead makes the hard tasks fail. Sending regulated data to a public endpoint causes compliance risks.

Engineering teams have started building their own answers. Uber’s engineering blog reported in August that changing models for its automated code review system improved quality scores while cutting cost per pull request, part of a program that brought cost per thousand model requests down about 34% from its peak. Sword Health built an internal proxy between its tools and providers and credits part of a 94% cost reduction to routing work to cheaper, self-hosted models.

The large clouds have productized a version of this. AWS made Bedrock Intelligent Prompt Routing generally available in April 2025 and says it can reduce costs by up to 30%. Microsoft’s Foundry Model Router, introduced at Build 2025, selects a model per prompt. Google Cloud added model routing to its API Gateway in August 2026. Each of these answers the question of what model to run. Yet none of them decides which country a request runs in, or whether it runs on hardware outside a hyperscale region.

GoodVision AI’s Smart Routing Engine scores every request on four things: complexity, latency requirement, cost and data sensitivity. It then picks both the model and the environment, public cloud for general work and dedicated or local compute for sensitive and latency-critical work, running on NVIDIA’s inference software. Wang describes the rule as the right model for the right task. The platform entered limited preview with invited enterprise customers in February 2026, the company began selling inference services the following month, and a broader release is still ahead of it.

Two deployments show the range. A gaming studio that had designed an in-game agent system, with no compute to run it and a team that wanted to work on gameplay, built it on the platform; GoodVision AI reports token costs down by up to 35% and development effort down by up to 60%. A biotechnology customer whose customer data onboarding had become a manual bottleneck automated the ingestion step and reports up to 40% lower token costs and 70% less manual work. Research the company commissioned from Frost & Sullivan models a hybrid setup, with sensitive and high-frequency work kept local, using roughly 70% fewer tokens than sending everything to one public API.

“Every request carries three constraints,” Wang says. “What it costs, how fast it has to come back, and where the data is allowed to go. A person can weigh those for one request. Nobody can do it for a million a day.”

Routing only helps if the compute is there

Placement runs into a physical limit. The International Energy Agency expects data center electricity use to more than double, to around 945 TWh by 2030, slightly more than Japan’s total consumption today. EPRI projects US data center consumption between 383 and 793 TWh by 2030, two to four times the 2024 level.

New capacity is slow to connect. JLL reported in January that the average wait for a grid connection in primary data center markets now exceeds four years. In the European Union, the IEA puts connection queues at two to ten years depending on the country, and at seven to ten years in the Frankfurt, London, Amsterdam, Paris and Dublin hubs.

Governments are funding capacity of their own. South Korea is building a national AI computing center budgeted at up to 2 trillion won, and in July the European Commission opened a call for up to seven AI gigafactories supported by as much as 10 billion euros in public funding.

A router can decide that a request belongs in Fukushima or Dallas rather than a hyperscale region. The decision is worth nothing without capacity there to take it.

GoodVision AI has signed non-binding agreements for three sites under the program it calls AI Factory: Fukushima in Japan, Texas, and Gimpo in South Korea. Each is designed to be compact, high-density and liquid-cooled, sized to start at between 1.5MW and 5MW and expand from there, and placed near users and inside the jurisdictions their data belongs to rather than at hyperscale. Initial deployment is targeted for this year. The company says a site can be running within 30 days of equipment installation. The routing platform currently runs on capacity bought from the major cloud providers, the same constraint most of its customers are under.

“People ask why a routing company would build compute sites,” Wang says. “Routing is a decision. A decision is worth nothing if there is no capacity where the work needs to run. The two halves have to be solved together.”

Someone still has to run it

The third problem is that an AI application in production never stops changing, and somebody has to absorb those changes. It gets the least attention because it only begins after launch. A provider retires the model version an application was built on, and prompts tuned for the old one behave differently on its replacement. Prices and rate limits change, and the economics of a feature change with them. A product gets attention and traffic jumps tenfold on a Tuesday. A new contract says a customer’s data cannot be processed outside the EU, so a workload has to move.

That work falls to whoever already runs the company’s cloud accounts, which is where GoodVision AI started. Its cloud team handles procurement, migration, performance tuning and daily operations across the major providers, work the company says has cut customers’ cloud bills, as distinct from their token costs, by 30% to 50% against standard public cloud pricing. It is the least novel of its three lines and, for most enterprise buyers, the one that decides whether the other two reach production at all.

“Inference is not a project with an end date,” Wang says. “It runs every day and it changes every month. Most teams underestimate that part.”

Turning inference demand into revenue

Cloud services remain the foundation of GoodVision AI’s business. According to the Form S-4 filed with the U.S. Securities and Exchange Commission, revenue was approximately $3.6 million in fiscal 2024, $7.7 million in fiscal 2025, and $24 million for the nine months ended June 30, 2026.

The next stage of growth depends in part on bringing AI Factories into operation. Management projections disclosed in the S-4 contemplate approximately 17.5 MW of deployed AI Factory capacity in fiscal 2027, increasing to approximately 40 MW in fiscal 2028. Based in part on the anticipated deployment of this capacity, management projects AI Factory computing service revenue of approximately $64 million in fiscal 2027 and $224 million in fiscal 2028. These projections depend on site development, equipment procurement, financing and customer commitments.

Growth on that scale is usually assumed to mean selling more tokens. Wang argues the opposite:

“Customers are buying the outcome of a business task, not a number of tokens.” By eliminating unnecessary calls and retries, better routing can reduce token consumption while increasing the value delivered.

The Form S-4 for the proposed combination with Calisa became effective on September 11, and Calisa’s shareholders vote on October 8. No closing date is confirmed. Wang says the proceeds would go first to new AI Factory sites, to further work on the Smart Routing Engine, and to expand its sales and research teams.

The advantage is moving underneath the model

Any company can call the same models through the same APIs. What is harder to copy is everything underneath: deciding where each task runs, holding capacity to run it there, and keeping that arrangement working while models, prices and rules keep moving. That is where the advantage in enterprise AI is going.

Plenty of it remains unsolved. Judging whether a routing decision was the right one is still difficult, because quality is harder to measure than cost. Compute is being built more slowly than demand is growing, and operating across several providers and locations adds complexity that has to be paid for somewhere.

“Anyone can buy the model,” Wang says. “The pilots that survive are the ones that get the next decision right: where the work runs, what it costs to run there, and whether the capacity exists at all. That is the inference layer AI runs on.”

Learn more about GoodVision AI at goodvision.ai.

More in Cloud

All Cloud »