Google says its new Gemini 3.8 Flash model ‘works harder’ but might cost more
Google has expanded its developer portfolio with Gemini 3.8 Flash, an updated lightweight model designed for deeper reasoning and automated tool orchestration. Unlike its predecessor, this version conducts multiple intermediate thinking steps and executes tool calls iteratively before returning an answer. While the base rate per token remains stable at launch, the extended reasoning process consumes more tokens per query, which could increase overall inference expenses for high-volume enterprise production workflows.
What this means for you
Engineering teams should closely monitor total token usage during extended reasoning cycles to prevent unexpected operational cost increases.
Put this to work
Try the tools in this story inside the AI Tool Lab, browse roles in the AI Job Portal, test your skills with Aptitude Tests, or read our in-depth guides.
Keep reading
More AI news
Anthropic CEO says it’s time to pump the brakes on AI
Anthropic CEO Dario Amodei has proposed a strategic shift in AI development, advocating for a measured pace rather than an unchecked race for intelligence. By inviting independent third-party assessors like METR to audit their systems, the company aims to establish a new standard for transparency and safety. This move signals a significant departure from the 'growth-at-all-costs' mentality, suggesting that the industry must prioritize verifiable safety protocols before deploying more powerful models to the public.
Read the briefAnthropic CEO outlines plan to ‘pace the frontier’
Anthropic CEO Dario Amodei has proposed a strategic shift aimed at decelerating the breakneck speed of artificial intelligence development. By advocating for a more deliberate 'pacing' of innovation, the company suggests that industry leaders should prioritize safety evaluations and societal impact assessments over mere technical capability benchmarks. This shift marks a significant departure from the competitive race between major AI labs, highlighting growing concerns that rapid deployment without adequate guardrails could pose existential or systemic risks to global infrastructure.
Read the briefY Combinator’s Garry Tan wants U.S. open-weight AI labs to ‘distill’ frontier models, too
Y Combinator leader Garry Tan is advocating for a shift in how leading artificial intelligence developers share their technology. He contends that since foundational models are built upon vast repositories of human-generated information, the resulting capabilities should be accessible as a public benefit. By encouraging labs to create smaller, distilled versions of frontier models, Tan hopes to democratize access to high-performance AI, moving away from closed-off ecosystems toward a more equitable distribution of innovative tools.
Read the brief