Unlock exclusive introductory pricing for the newly launched Gemini 3.5 Flash.
G

Gemini 3.5 Flash Lite

Ввод:$0.24/M
Вывод:$2.016/M
Дата выпуска:Jul 21, 2026

Gemini 3.5 Flash-Lite is a low-latency, cost-effective multimodal model optimized for high-throughput, low-cost execution for subagent tasks and document parsing. The model supports text, image, video, audio, and PDF inputs, and is designed for high-volume agentic workflows, simple data extraction, and applications where latency and API cost are the primary constraints

Новый
Коммерческое использование

Playground для Gemini 3.5 Flash Lite

Изучите Playground Gemini 3.5 Flash Lite — интерактивную среду для тестирования моделей и выполнения запросов в реальном времени. Попробуйте промпты, настройте параметры и итерируйте мгновенно, чтобы ускорить разработку и проверить варианты использования.

Technical Specifications of Gemini 3.5 Flash-Lite

ItemGemini 3.5 Flash-Lite
ProviderGoogle DeepMind
Model IDgemini-3.5-flash-lite
Model familyGemini 3.5
AvailabilityGeneral Availability (GA)
Input typesText, Image, Video, Audio, PDF
Output typesText
Context window1,048,576 tokens
Maximum output65,536 tokens
ThinkingSupported (default: minimal; also medium and high)
Function callingYes
Code executionYes
File SearchYes
URL ContextYes
Search GroundingYes
Structured OutputYes
Computer UseNot supported
CachingSupported
Primary focusHigh-throughput, low-latency, low-cost inference

What is Gemini 3.5 Flash-Lite?

Gemini 3.5 Flash-Lite is Google's fastest and most cost-effective model in the Gemini 3.5 family. It is optimized for workloads where latency, throughput, and API cost are more important than maximum reasoning performance. Typical applications include document parsing, structured data extraction, routing, lightweight coding, and autonomous subagents operating at scale. Google positions it as the recommended upgrade path from Gemini 3.1 Flash-Lite and Gemini 2.5 Flash for production deployments.

Main Features of Gemini 3.5 Flash-Lite

  • Optimized for high-volume, low-cost inference.
  • Supports a 1 million-token context window for long documents and repositories.
  • Native multimodal inputs including text, images, video, audio, and PDF.
  • Supports Function Calling, Code Execution, File Search, URL Context, Search Grounding, and Structured Outputs.
  • Configurable thinking levels (minimal, medium, high) to balance speed and reasoning quality.
  • Significantly improves coding, reasoning, and agentic workflows compared with Gemini 3.1 Flash-Lite while maintaining very low latency.

Benchmark Performance of Gemini 3.5 Flash-Lite

Google states that Gemini 3.5 Flash-Lite substantially outperforms previous Flash-Lite generations in coding, document understanding, and agentic workflows while remaining the lowest-cost model in the Gemini 3.5 lineup. Gemini 3.5 Flash-Lite scored 54% in the Terminal-Bench 2.1 test, a significant improvement over its predecessor's 31%.

Its strengths lie in classification, extraction, chat replies, and simple retrieval-enhanced answers. These are precisely where fast, low-cost models can excel.

Gemini 3.5 Flash-Lite vs Gemini 3.6 Flash vs Gemini 3.5 Flash

AspectGemini 3.5 Flash-LiteGemini 3.6 FlashGemini 3.5 Flash
PositioningFastest & cheapest for high-volume everyday tasks (e.g., document processing, search)Current flagship Flash: Best balance of intelligence, efficiency & agentic performanceStrong agentic/coding model (now largely superseded by 3.6)
Best ForHigh-throughput, low-cost workflowsCoding, multimodal, complex agents, knowledge workGeneral agentic & coding tasks
Intelligence / PerformanceGood (outperforms older Lites; competitive on many agentic tasks)Highest among the three (noticeable gains in coding & agentic)Strong (near-frontier for Flash series)
Key Benchmarks- Terminal-Bench: ~54%- Strong on document & high-volume tasks- DeepSWE: 49% (vs 37%)- MLE-Bench: 63.9% (vs 49.7%)- OSWorld: 83% (vs 78.4%)- Terminal-Bench: 76.2%- Lower than 3.6 on most coding/agentic
SpeedFastest (~350 output tokens/sec)Very fast (~280 t/s)Fast
EfficiencyExcellent for volumeBest (17% fewer output tokens on average; up to 65% on coding)Good
MultimodalText + Image + Video + AudioText + Image + Video + Audio (stronger reasoning)Text + Image + Video + Audio
Context Window~1M tokens~1M tokens~1M tokens
Max Output Tokens~64k~64k~64k
Pricing (per 1M tokens)Cheapest: ~$0.30 input / $2.50 output$1.50 input / $7.50 output$1.50 input / $9.00 output
Effective CostLowest per task for simple workLower than 3.5 due to efficiency gainsHigher than 3.6

Summary Recommendations

  • Pick 3.5 Flash-Lite — when you need maximum speed and minimum cost for high-volume or simple tasks.
  • Pick 3.6 Flash — for most users and developers (best overall choice right now).
  • Pick 3.5 Flash — only if you have existing workflows tied to it (plan to upgrade to 3.6).

Limitations of Gemini 3.5 Flash-Lite

  • Optimized for efficiency rather than frontier-level reasoning.
  • Computer Use is not supported.
  • Native image and audio generation are unavailable.
  • Fine-tuning is not currently supported.
  • Complex multi-step reasoning tasks are generally better suited to Gemini 3.5 Flash or Gemini 3.6 Flash.

ЧАВО

Цены для Gemini 3.5 Flash Lite

Изучите конкурентоспособные цены на Gemini 3.5 Flash Lite, разработанные для различных бюджетов и потребностей использования. Наши гибкие планы гарантируют, что вы платите только за то, что используете, что упрощает масштабирование по мере роста ваших требований. Узнайте, как Gemini 3.5 Flash Lite может улучшить ваши проекты, сохраняя при этом управляемые расходы.

ModelComet Price (USD / M Tokens)Official Price (USD / M Tokens)Discount
gemini-3.5-flash-lite
Ввод:$0.24/M
Вывод:$2.016/M
Ввод:$0.3/M
Вывод:$2.52/M
-20%

Пример кода и API для Gemini 3.5 Flash Lite

Получите доступ к исчерпывающим примерам кода и ресурсам API для Gemini 3.5 Flash Lite, чтобы упростить процесс интеграции. Наша подробная документация предоставляет пошаговые инструкции, помогая вам использовать весь потенциал Gemini 3.5 Flash Lite в ваших проектах.

#!/bin/bash

curl "https://api.cometapi.com/v1beta/models/gemini-3.5-flash-lite:generateContent" \
  -H "x-goog-api-key: $COMETAPI_KEY" \
  -H 'Content-Type: application/json' \
  -X POST \
  -d '{
    "contents": [
      {
        "parts": [
          {
            "text": "Write a three.js script that renders an interactive 3D robot."
          }
        ]
      }
    ],
    "generationConfig": {
      "maxOutputTokens": 8192
    }
  }'

cURL Code Example

#!/bin/bash

curl "https://api.cometapi.com/v1beta/models/gemini-3.5-flash-lite:generateContent" \
  -H "x-goog-api-key: $COMETAPI_KEY" \
  -H 'Content-Type: application/json' \
  -X POST \
  -d '{
    "contents": [
      {
        "parts": [
          {
            "text": "Write a three.js script that renders an interactive 3D robot."
          }
        ]
      }
    ],
    "generationConfig": {
      "maxOutputTokens": 8192
    }
  }'

Python Code Example

import os

from google import genai
from google.genai import types


client = genai.Client(
    api_key=os.environ["COMETAPI_KEY"],
    http_options={
        "api_version": "v1beta",
        "base_url": "https://api.cometapi.com",
        "retry_options": {"attempts": 1},
    },
)

response = client.models.generate_content(
    model="gemini-3.5-flash-lite",
    contents="Write a three.js script that renders an interactive 3D robot.",
    config=types.GenerateContentConfig(max_output_tokens=8192),
)

print(f"Response ID: {response.response_id}")
print(f"Model Version: {response.model_version}")
print(f"Finish Reason: {response.candidates[0].finish_reason}")
print(response.text)

JavaScript Code Example

import { GoogleGenAI } from "@google/genai";


const ai = new GoogleGenAI({
  apiKey: process.env.COMETAPI_KEY,
  httpOptions: {
    apiVersion: "v1beta",
    baseUrl: "https://api.cometapi.com",
    retryOptions: { attempts: 1 },
  },
});

const response = await ai.models.generateContent({
  model: "gemini-3.5-flash-lite",
  contents: "Write a three.js script that renders an interactive 3D robot.",
  config: {
    maxOutputTokens: 8192,
  },
});

console.log(`Response ID: ${response.responseId}`);
console.log(`Model Version: ${response.modelVersion}`);
console.log(`Finish Reason: ${response.candidates[0].finishReason}`);
console.log(response.text);

Uptime

Процент успешных запросов за последние 30 дней, отражающий надёжность каждого поставщика моделей. CometAPI круглосуточно отслеживает всех подключённых поставщиков в режиме реального времени.

RespondLIVE
1655msAvg. Response
UptimeLIVE
100.0%Avg. Uptime

Версии Gemini 3.5 Flash Lite

Причина наличия нескольких снимков Gemini 3.5 Flash Lite может включать такие потенциальные факторы, как: изменения в выходных данных после обновлений, требующие сохранения старых снимков для обеспечения согласованности; предоставление разработчикам переходного периода для адаптации и миграции; а также наличие разных снимков, соответствующих глобальным или региональным конечным точкам для оптимизации пользовательского опыта. Для получения подробной информации о различиях между версиями обратитесь к официальной документации.

Version
gemini-3.5-flash-lite