← Back to Blog
Announcement · 3 min read

Gemini 3.8 Flash Is Now in Ayeto: A Fast Model for Agents

Team AYETO ·
Gemini 3.8 Flash Is Now in Ayeto: A Fast Model for Agents

Long-horizon performance without premium-model pricing

We have added Gemini 3.8 Flash to Ayeto. Google built this workhorse model for long-horizon software engineering, autonomous agents, and complex enterprise workflows. It builds on Gemini 3.7 Flash but focuses more strongly on tasks that require repeated tool use, result verification, and several coordinated steps.

Google released the model on September 2, 2026. A separate Gemini 3.8 Flash Cyber variant is available only to trusted defenders through the Fairwind Program. The model available for general workflows in Ayeto is the standard Gemini 3.8 Flash.

What is Gemini 3.8 Flash?

Gemini 3.8 Flash is a multimodal model optimized for fast, cost-efficient agentic workflows. It accepts text, images, video, audio, and PDFs and produces text. It has a 1,048,576-token context window and supports up to 65,536 output tokens.

Its capabilities include function calling, structured outputs, code execution, file search, Google Search grounding, URL context, and computer use in Preview. Reasoning effort can be set to low, medium, or high; minimal is not supported.

What changes from Gemini 3.7 Flash

Google describes 3.8 Flash as a more diligent model. On complex tasks, it can perform additional reasoning steps and call tools iteratively to verify and refine the result. This improves long-horizon execution, although higher effort may also increase token usage.

The model scores 54.9% on HLE-Verified. Google also reports a substantial improvement on DeepSWE v1.1 for long-horizon software engineering and stronger results on financial and legal agent benchmarks. Benchmarks are useful orientation, but your own workflow remains the meaningful test.

When to choose Gemini 3.8 Flash in Ayeto

The model is a strong fit for:

  • multi-file code fixes and longer development tasks,
  • assistants that repeatedly use search, APIs, or other tools,
  • analysis of large PDFs, videos, and technical documentation,
  • reports assembled from several sources,
  • automation where performance, speed, and cost must stay balanced.

A smaller model may be more efficient for short classification and simple transformations. For the hardest tasks where maximum capability matters more than cost, compare Flash with premium frontier models.

How the cost compares

Through the end of 2026, Gemini 3.8 Flash has the same introductory API price as Gemini 3.7 Flash. It therefore adds capability without increasing the unit rate. Compared with GPT-6 Astra, standard API pricing makes it approximately 13× less expensive for both input and output.

That does not guarantee a completed task will cost 13× less. Gemini 3.8 Flash may use more reasoning steps and thinking tokens on complex work. Compare total consumption for the finished result rather than the rate per million tokens alone.

Submit the same prompt to several models through Ayeto Model Comparison to compare output, latency, and credit usage side by side.

Try Gemini 3.8 Flash on a real workflow

Select Gemini 3.8 Flash in your assistant settings and give it a task with several stages: analyze source material, use a tool, and verify the result. Gemini 3.8 Flash is now available in Ayeto, where you can compare it without moving your workflow to another provider.

Sources