Jump to Content
AI & Machine Learning

What sports cars can teach us about optimizing AI token spend (Q&A)

August 26, 2026
https://storage.googleapis.com/gweb-cloudblog-publish/images/10-companies-making-ai-work-across-industr.max-2600x2600.jpg
Mike Clark

Director of Product Management, Gemini Enterprise Agent Platform

Natasha Balasubramanian

Product Marketing Manager, Google Cloud

Try Gemini Enterprise today

The front door to AI in the workplace

Try now

Many of us who work in the AI space have heard the trend “tokenmaxxing.” Broadly, this refers to the belief that simply increasing token consumption will somehow translate to better AI.

Mike Clark, Director of Product Management for Gemini Enterprise Agent Platform, thinks differently. To him, unlocking AI’s true potential relies on a strategy of maxing out efficiency, rather than tokens.

With a career spent scaling some of the internet's most recognizable platforms, from early social web pioneers like Photobucket and VSCO to foundational generative AI infrastructure, Mike is now focused on shaping the future of how we work.

We sat down with Mike to talk about why the “more is always better” hype in AI is failing enterprises, and how Gemini Enterprise is helping leaders take control of their AI spend. What follows is an edited version of our conversation.

Q: Can you tell me about your background and your professional journey, and what brought you to Google?

Mike: My career began at Philips Electronics before I transitioned into the startup world. Over the years, I helped scale Photobucket to hundreds of millions of users, built Safety Web to help parents navigate their kids' online safety, and led engineering at VSCO.

More recently, I focused on privacy infrastructure before establishing generative AI data practices for a major foundation model launch, eventually spending my time with research building AI agents.

I ultimately joined Google Cloud because I value strong leadership and wanted to collaborate with the best talent in the industry. The AI work we are doing right now is actively shaping the future of how people work and communicate, and I want to be at the forefront of this shift, helping to make a positive, scalable impact on the world.

Q: You’ve compared managing AI compute costs to driving a sports car with an "S" (Sport) and "E" (Economy) button. Can you break down that analogy for us?

Mike: Think of it like a dashboard on a sports car. When you toggle between "Economy" or "Sport" modes, you can watch the mileage gauge on your dash change in real time – say from from 30 miles per gallon down to 12. You definitely feel the extra horsepower in Sport mode, but you also understand the tradeoff: you're maybe only going to get 100 miles out of this tank instead of 300, but if you have a mountain to climb, or need to accelerate quickly, that extra performance is exactly what you're trading your fuel for.

In an enterprise AI environment, leaders need that same dashboard. They need to see exactly what they are trading their "fuel" – or token budget – for, and have the controls to toggle modes based on the task at hand.

Q: What does pressing that "Economy" button look like for an enterprise running agents?

Mike: It starts with recognizing that not every agent task requires real-time, low-latency processing.

For instance, if you have a user interacting with a customer-facing chatbot, they need a response in seconds. That’s Sport mode. But if your agents are running background tasks – like a massive nightly data reconciliation or a deep code-repository refactoring sweep – nobody is sitting there waiting on a loading screen.

For those background processes, you can press the "Economy" button. This means deferring those agent workloads to off-peak hours, or setting them to trickle-process over a wider window – like over multiple days at a time. You still get to use the best models, but at a fraction of the price without burning through a full tank of gas.

Another lever is token optimization at the orchestration level. A great example of this is how you might use a faster, lower-capability model for the majority of a task, but consult the most capable models only at specific times when needed. Again, you get the best outcome, but with a highly optimized token footprint.

Q: Right now, there is a massive industry trend of "tokenmaxxing." Why do you think that's the wrong approach?

Mike: Some companies are making some really broad assumptions that just aren't true, like the idea that they can simply replace an employee with more tokens. It's almost like they get hooked on the idea that AI is going to solve absolutely everything, and the way they try to achieve that is by just consuming more and more tokens.

But that’s not how it plays out. You don't get the most out of AI by just throwing more tokens at a task; you do it by focusing on outcomes.

Q: If "tokenmaxxing" isn't the answer, what should enterprises be doing instead?

Mike: Following the same analogy and spirit, I would advocate for "efficiencymaxxing” instead. It’s about getting the highest-quality business outcomes out of every dollar of what you're spending. The clearer you are able to express your desired outcome to a model, the better the plan and the results you'll get.

To build that foundation, I like to frame enterprise cost management around three basic goals: predicting what you're going to spend, controlling what you're actually spending, and understanding your bill after the fact to see the value you got.

Companies are realizing the real magic happens when you combine the power of the employees you already have with optimized agentic workflows – making your people significantly more capable, while freeing them to focus on the things that actually drive your business.

Q: How can Gemini Enterprise help?

Mike: On Gemini Enterprise, we’re addressing these challenges in two ways. First, we’re building cost management capabilities directly into the core platform architecture to help teams plan, enforce, and understand their spend.

Second, we’re building those efficiency choices we talked about directly into the products – whether that is workload deferral options or intelligent, orchestration-level token routing. This gives organizations the freedom to pursue greater efficiency without ever having to sacrifice output quality.

Q: What advice do you give to companies trying to build better cost transparency and control into their agent rollouts today?

Mike: I always tell them to start by defining a clear outcome for your business. What is the core thing you are trying to accomplish? The better you articulate that, the better the agents will perform.

First, plan your budget upfront to balance token spend and infrastructure. Next, implement tight controls to enforce boundaries, like spend caps, so you don't run away on a process that isn't driving value. Finally, always audit your agent to understand the actual outcomes they delivered.

When you treat agent planning with that level of rigor, you aren't just making existing tasks run a little faster. You're clearing away the undifferentiated toil, giving your people the space to do their best work, and fundamentally changing what your business is capable of.

Posted in