This website uses cookies

Read our Privacy policy and Terms of use for more information.

The landscape of frontier AI models is undergoing a significant structural shift. Google recently announced the release of three new, highly specialised models within its Gemini family: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. Concurrently, the release of the much-anticipated flagship model, Gemini 3.5 Pro, remains delayed while the company begins pre-training for Gemini 4.

This development marks a clear departure from the pursuit of a single, omnipotent model. Instead, we are seeing the fragmentation of the frontier into task-priced specialists. For project delivery teams integrating AI into their workflows, this necessitates a more sophisticated approach to model routing and cost management.

The Rise of the Specialists

The new lineup is designed to address specific operational bottlenecks.

According to Google's official announcement, the Gemini team articulated the strategic rationale: "Developers and customers building production AI agents need higher token efficiency, lower latency, and more reliable performance. Our Flash series of models is built to meet the sweet spot of efficiency and quality to enable scaling agentic workflows."


Gemini 3.6 Flash is positioned as the "agent workhorse." It offers a 17% reduction in output token usage compared to its predecessor, 3.5 Flash, while demonstrating improved performance in coding and complex knowledge work. Priced at $1.50 per million input tokens and $7.50 per million output tokens, it is built for sustained, multi-step agentic workflows.

For high-volume, low-latency tasks, Google introduced Gemini 3.5 Flash-Lite. Operating at an impressive 350 output tokens per second, it is the fastest model in the 3.5 series. At just $0.30 per million input tokens, it is highly cost-effective for tasks like agentic search and rapid document processing.

Finally, Gemini 3.5 Flash Cyber represents a highly targeted approach. This model is specialised for vetted-partner vulnerability hunting and is paired directly with Google's CodeMender code security agent, illustrating how models are increasingly being coupled with specific software infrastructure to solve niche problems.

Rethinking Workload Routing

The era of defaulting to the largest, most capable model for every task is over. The cost implications of using a flagship model for routine data extraction or basic drafting are simply too high when specialised, cheaper models can perform the same tasks with equal or greater efficiency.

Delivery teams must now implement intelligent routing systems. Complex reasoning, strategic planning, or nuanced stakeholder communication might still require the capabilities of a flagship model. However, high-volume tasks—such as processing thousands of site photos, scanning supply chain documents, or running preliminary code checks—should be routed to models like 3.5 Flash-Lite to optimise both speed and expenditure.

Google's pricing structure for these models reflects the market's recognition of this segmentation. Gemini 3.6 Flash, at $1.50 per million input tokens, sits at a midpoint between the ultra-cheap 3.5 Flash-Lite and what we would expect from a flagship model. This pricing architecture incentivises developers to think strategically about task routing. A firm processing 100 million tokens per month could save tens of thousands of pounds annually by routing appropriate workloads to the lower-cost alternatives.

The Implications of the Pro Delay

The delay of Gemini 3.5 Pro, while Google focuses on the pre-training of Gemini 4, is also telling. It suggests that the marginal gains achieved by iterating on current architectures are becoming harder to realise, prompting labs to look toward the next major paradigm shift.

This pause in flagship releases provides a crucial window for organisations to optimise their current implementations. Rather than constantly chasing the newest, largest model, teams should focus on building robust, modular architectures that can seamlessly swap in specialised models as they become available. The competitive advantage will belong to those who can orchestrate these diverse models effectively, rather than those who simply pay for the most expensive API access.

The fragmentation we are witnessing is not a temporary phenomenon; it is the beginning of a more mature market structure. As AI capabilities become increasingly commoditised, differentiation will shift from raw model capability to the sophistication of the systems built around those models. Firms that can implement intelligent routing, cost optimisation, and seamless model orchestration will outperform those that treat AI as a black box.

Strategic Implications for Project Delivery

We must adapt our digital strategies to accommodate this fragmented landscape. When designing automated workflows for cost estimation, risk analysis, or design review, the architecture must dictate which model handles which component of the task.

For instance, a system reviewing a complex engineering specification might use a fast, lightweight model to extract the text and identify key clauses before passing only the highly technical or ambiguous sections to a more capable, expensive model for deep analysis. This modular approach not only controls costs but also significantly reduces latency, making real-time AI assistance a viable reality on the project site.

Takeaway

Embrace modular architecture: Systems must be designed to route tasks dynamically to the most appropriate model based on the specific requirements for speed, cost, and reasoning capability.

Stop defaulting to the flagship: Using the largest, most expensive model for routine tasks is financially inefficient; leverage specialised models like Flash-Lite for high-volume data processing.

Optimise for token efficiency: The introduction of models like 3.6 Flash highlights the importance of token efficiency; workflows should be designed to minimise unnecessary output to control costs.

Prepare for deeper specialisation: As models become increasingly tailored to specific domains (like Flash Cyber), firms should actively seek out models that align with their particular operational niches.

Get ahead of the curve

How are you managing the cost and complexity of your AI deployments in a fragmenting market? Subscribe to the Project Flux newsletter for strategies on building cost-effective, multi-model workflows in project delivery.

Links and Stuff

All content reflects our personal views and is not intended as professional advice or to represent any organisation.

1  

Reply

Avatar

or to participate

Keep Reading