Seedance 2.0 Mini & Fast API với mức giá thấp nhất toàn cầu — giảm đến 68% so với giá chính thức

How Lindy Cut Its Agent Inference Costs by 90% on Atlas Cloud

Lindy moved its AI employees to DeepSeek v4 Flash on Atlas Cloud and cut agent inference costs 90%. See how prompt caching held the savings at 10x scale.

How Lindy Cut Its Agent Inference Costs by 90% on Atlas Cloud

Lindy built the AI employee. It moved the fleet onto open weights served by Atlas Cloud, cut its inference bill by about 90%, and did it without shipping a worse product.

~90% lower inference cost · 60% of input tokens served from cache · 3,000+ requests per minute sustained · 10x+ traffic growth, nothing rebuilt · 24 models from 11 labs on one key

Lindy runs one of the most demanding agent workloads in production, and it runs on Atlas Cloud. Here is what that changed:

  • Because Atlas prices the repeated call at cache rates, Lindy runs agents that check their work instead of guessing. Six of every ten input tokens it sends are served from cache, so the prefix an agent re-reads a dozen times a task costs almost nothing.
  • Because Atlas provisions for the workload, Lindy grew its volume more than tenfold by sending more traffic, not by rebuilding anything, and Atlas holds above 3,000 requests per minute for an hour at a stretch.
  • Because Atlas carries the whole catalog on one key, Lindy can swap to a new model in an afternoon. It has run 24 models from 11 labs through a single integration.
  • Because Atlas serves it on a named, SOC 2 Certified contract, Lindy can put an open-weight model in front of its customers' data and stand behind who is running it.

The 90% is Lindy's headline. The reason it holds, and the reason the next migration will be easier than this one, is the platform underneath.

For Lindy, the bill is the business

Lindy builds AI employees. A Lindy Teammate joins a company the way a new hire does: it takes requests in Slack, connects to the tools the team already runs, sits in meetings, and keeps what it learns so the next request starts further along than the last. Nobody writes an automation; they delegate, and the agent does the work.

That design sets a hard infrastructure bill. An AI employee that reads a thread, checks a calendar, looks up a record and drafts a reply has made a dozen model calls before anyone sees a word, and because a Teammate serves a whole team, volume grows with headcount. At that shape, the price of the model behind the product determines business viability.

[PUBLIC] "Lindy's pricing only works if inference keeps getting cheaper."

— Bruno Škvorc, Staff Software Engineer, Lindy

So Lindy did the thing that finance math demands. It moved most of its managed-agent traffic off Claude, Sonnet and Gemini onto DeepSeek v4 Flash running on Atlas Cloud, and inference costs on the migrated routes fell by about 90%. Changing the model name took an afternoon. Making it stick, at production volume, without the product getting worse, is why they chose Atlas Cloud.

Why the 90% holds: the eleventh call is nearly free

Priced by the token, an agent pays full freight for the same prefix a dozen times per task, and the bill grows with how carefully it thinks. That tax is what keeps agent products shallow: one pass instead of three, because the third costs as much as the first. It is also why a naive model swap saves less than the sticker suggests, as the repeated prefix quietly refills the bill.

Atlas removes the tax where it is heaviest. Six of every ten of Lindy's input tokens are recognized rather than reprocessed, at a rate contracted against its real volume, so the prefix that dominates every agent call is close to free. That is what turns a model switch into a durable 90% rather than a number that erodes as usage climbs, and it is what lets an agent on Atlas run three passes where the same agent, priced by the token, would run one.

[PROPOSED — your call] "Agent workloads reuse a lot of context across many model calls. Atlas’s caching means we don’t pay full price for that same context every time, which is a big part of why the savings held up at scale.."

— Ian McGregor, Head of Engineering, Lindy

Capacity that was there before Lindy needed it

Agent traffic has no quiet overnight window and no launch day to plan around. Work arrives while teams are working and does not stop while you scale. What matters is not the spike a buffer can absorb but the rate a provider can hold. Lindy grew its volume more than tenfold by sending more traffic, not by rebuilding anything, and Atlas held above 3,000 requests per minute for an hour at a stretch. Capacity was ahead of the workload, so scaling was a business decision and never an infrastructure project.

Support that ships product, not tickets

At this volume the difference between providers is less the dashboard than who answers when something looks wrong. Lindy's engineers and Atlas's inference engineers share a channel, and answers come back the same day from people who can act on them. Some of those answers become product changes: Lindy asked for a way to transfer ownership of a team account, Atlas did not support it at the time, and it shipped, with Atlas's engineers moving the account over themselves.

A provider you can name

Moving to an open-weight model removes the party that used to answer for how the model was run. The questions a lab's brand used to settle now point at the provider: who is serving this, under what controls, and what happens to the data passing through. Atlas answers those by name rather than by router, on a SOC 2 Certified contract, with client data neither stored nor used for training. That matters more for an AI employee than for a chat product, because a Teammate reads the Slack threads, calendars and records of the company it works for. The infrastructure question sits directly under the customer-trust question, and Atlas is the answer to both.

Not locked to DeepSeek, locked to nothing

The migration committed Lindy to a strategy, not to a model. DeepSeek v4 Flash won the workloads it was tested on and keeps the position for as long as it is the best price at adequate quality. The next winner will come from a different lab under a different license, and because Atlas Cloud provides access to all models on one API, trying it costs an afternoon rather than a procurement cycle. In one day of testing, Lindy ran 47 requests through twelve models it had never used before, across seven labs and three modalities, all on the key it already had. Three became production workloads. The freedom to always run the best model through Atlas Cloud gives Lindy true operational business fluidity.

If you are running agents on a marketplace, move them

Lindy walked this exact path. It first met DeepSeek on OpenRouter, ran the model through its evals there, and proved it was worth a migration. In that same testing it found the thing that decided where production would run.

[PUBLIC] "We also tested the same model on different inference providers. Annoyingly, the provider mattered. The same nominal model could score differently depending on who served it."

— Bruno Škvorc, Staff Software Engineer, Lindy

A marketplace is built to help you shop, not to run your product. Send production traffic through a router and it goes to whichever provider has spare capacity, so you do not choose who serves your model and you cannot see who did. The same weights on different machines return different numbers, from quantization or a shortcut in someone's serving stack, and those numbers reach your users before they reach your dashboard. Every router hop silently re-rolls the quality of the product your customers are paying for. You cannot debug it, because you cannot see who served the call. You cannot fix it, because you do not have full control of the routing. That is your reputation, decided by a coin you never get to flip.

So Lindy did not run production on OpenRouter. When the traffic went live, it went onto a direct contract with Atlas: one serving stack, the same one every request, tuned for the model and held to a contract, with the caching and capacity a live agent workload needs and the full catalog on the same key. If your agents are in production and still running through a router, you are shipping a product you cannot hold steady, and you will not know it until a customer does. Move them, the way Lindy did.

Talk to us about your workload and we'll tell you what it should cost, or browse the catalog.

About Lindy

Lindy builds AI employees. Lindy Teammate, launched in August 2026, works alongside a human team: it takes requests in Slack, connects to the tools a company already runs, joins meetings, and builds up the team's context so every request starts further along than the last. Instead of asking people to build and maintain automations, Lindy asks them to delegate. Lindy was founded by Flo Crivello and is based in San Francisco.

About Atlas Cloud

Atlas Cloud is a unified full-modal AI inference platform: 400+ models across video, image, language, and audio, through one API key, one endpoint, and one billing account, OpenAI-compatible for language models. SOC 2 Certified.

Mô hình mới nhất

Một API cho mọi AI đa phương tiện.

Khám phá tất cả mô hình