Unlock exclusive introductory pricing for the newly launched Gemini 3.5 Flash.
G

Gemini 3.5 Flash Lite

Entrรฉe:$0.24/M
Sortie:$2.016/M
Publiรฉ:Jul 21, 2026

Gemini 3.5 Flash-Lite is a low-latency, cost-effective multimodal model optimized for high-throughput, low-cost execution for subagent tasks and document parsing. The model supports text, image, video, audio, and PDF inputs, and is designed for high-volume agentic workflows, simple data extraction, and applications where latency and API cost are the primary constraints

Nouveau
Usage commercial

Playground pour Gemini 3.5 Flash Lite

Explorez le Playground de Gemini 3.5 Flash Lite โ€” un environnement interactif pour tester les modรจles et exรฉcuter des requรชtes en temps rรฉel. Essayez des invites, ajustez les paramรจtres et itรฉrez instantanรฉment pour accรฉlรฉrer le dรฉveloppement et valider les cas d'utilisation.

Technical Specifications of Gemini 3.5 Flash-Lite

ItemGemini 3.5 Flash-Lite
ProviderGoogle DeepMind
Model IDgemini-3.5-flash-lite
Model familyGemini 3.5
AvailabilityGeneral Availability (GA)
Input typesText, Image, Video, Audio, PDF
Output typesText
Context window1,048,576 tokens
Maximum output65,536 tokens
ThinkingSupported (default: minimal; also medium and high)
Function callingYes
Code executionYes
File SearchYes
URL ContextYes
Search GroundingYes
Structured OutputYes
Computer UseNot supported
CachingSupported
Primary focusHigh-throughput, low-latency, low-cost inference

What is Gemini 3.5 Flash-Lite?

Gemini 3.5 Flash-Lite is Google's fastest and most cost-effective model in the Gemini 3.5 family. It is optimized for workloads where latency, throughput, and API cost are more important than maximum reasoning performance. Typical applications include document parsing, structured data extraction, routing, lightweight coding, and autonomous subagents operating at scale. Google positions it as the recommended upgrade path from Gemini 3.1 Flash-Lite and Gemini 2.5 Flash for production deployments.

Main Features of Gemini 3.5 Flash-Lite

  • Optimized for high-volume, low-cost inference.
  • Supports a 1 million-token context window for long documents and repositories.
  • Native multimodal inputs including text, images, video, audio, and PDF.
  • Supports Function Calling, Code Execution, File Search, URL Context, Search Grounding, and Structured Outputs.
  • Configurable thinking levels (minimal, medium, high) to balance speed and reasoning quality.
  • Significantly improves coding, reasoning, and agentic workflows compared with Gemini 3.1 Flash-Lite while maintaining very low latency.

Benchmark Performance of Gemini 3.5 Flash-Lite

Google states that Gemini 3.5 Flash-Lite substantially outperforms previous Flash-Lite generations in coding, document understanding, and agentic workflows while remaining the lowest-cost model in the Gemini 3.5 lineup. Gemini 3.5 Flash-Lite scored 54% in the Terminal-Bench 2.1 test, a significant improvement over its predecessor's 31%.

Its strengths lie in classification, extraction, chat replies, and simple retrieval-enhanced answers. These are precisely where fast, low-cost models can excel.

Gemini 3.5 Flash-Lite vs Gemini 3.6 Flash vs Gemini 3.5 Flash

AspectGemini 3.5 Flash-LiteGemini 3.6 FlashGemini 3.5 Flash
PositioningFastest & cheapest for high-volume everyday tasks (e.g., document processing, search)Current flagship Flash: Best balance of intelligence, efficiency & agentic performanceStrong agentic/coding model (now largely superseded by 3.6)
Best ForHigh-throughput, low-cost workflowsCoding, multimodal, complex agents, knowledge workGeneral agentic & coding tasks
Intelligence / PerformanceGood (outperforms older Lites; competitive on many agentic tasks)Highest among the three (noticeable gains in coding & agentic)Strong (near-frontier for Flash series)
Key Benchmarks- Terminal-Bench: ~54%- Strong on document & high-volume tasks- DeepSWE: 49% (vs 37%)- MLE-Bench: 63.9% (vs 49.7%)- OSWorld: 83% (vs 78.4%)- Terminal-Bench: 76.2%- Lower than 3.6 on most coding/agentic
SpeedFastest (~350 output tokens/sec)Very fast (~280 t/s)Fast
EfficiencyExcellent for volumeBest (17% fewer output tokens on average; up to 65% on coding)Good
MultimodalText + Image + Video + AudioText + Image + Video + Audio (stronger reasoning)Text + Image + Video + Audio
Context Window~1M tokens~1M tokens~1M tokens
Max Output Tokens~64k~64k~64k
Pricing (per 1M tokens)Cheapest: ~$0.30 input / $2.50 output$1.50 input / $7.50 output$1.50 input / $9.00 output
Effective CostLowest per task for simple workLower than 3.5 due to efficiency gainsHigher than 3.6

Summary Recommendations

  • Pick 3.5 Flash-Lite โ€” when you need maximum speed and minimum cost for high-volume or simple tasks.
  • Pick 3.6 Flash โ€” for most users and developers (best overall choice right now).
  • Pick 3.5 Flash โ€” only if you have existing workflows tied to it (plan to upgrade to 3.6).

Limitations of Gemini 3.5 Flash-Lite

  • Optimized for efficiency rather than frontier-level reasoning.
  • Computer Use is not supported.
  • Native image and audio generation are unavailable.
  • Fine-tuning is not currently supported.
  • Complex multi-step reasoning tasks are generally better suited to Gemini 3.5 Flash or Gemini 3.6 Flash.

FAQ

Tarification pour Gemini 3.5 Flash Lite

Dรฉcouvrez des tarifs compรฉtitifs pour Gemini 3.5 Flash Lite, conรงus pour s'adapter ร  diffรฉrents budgets et besoins d'utilisation. Nos formules flexibles garantissent que vous ne payez que ce que vous utilisez, ce qui facilite l'adaptation ร  mesure que vos besoins รฉvoluent. Dรฉcouvrez comment Gemini 3.5 Flash Lite peut amรฉliorer vos projets tout en maรฎtrisant les coรปts.

ModelComet Price (USD / M Tokens)Official Price (USD / M Tokens)Discount
gemini-3.5-flash-lite
Entrรฉe:$0.24/M
Sortie:$2.016/M
Entrรฉe:$0.3/M
Sortie:$2.52/M
-20%

Exemple de code et API pour Gemini 3.5 Flash Lite

Accรฉdez ร  des exemples de code complets et aux ressources API pour Gemini 3.5 Flash Lite afin de simplifier votre processus d'intรฉgration. Notre documentation dรฉtaillรฉe fournit des instructions รฉtape par รฉtape pour vous aider ร  exploiter tout le potentiel de Gemini 3.5 Flash Lite dans vos projets.

#!/bin/bash

curl "https://api.cometapi.com/v1beta/models/gemini-3.5-flash-lite:generateContent" \
  -H "x-goog-api-key: $COMETAPI_KEY" \
  -H 'Content-Type: application/json' \
  -X POST \
  -d '{
    "contents": [
      {
        "parts": [
          {
            "text": "Write a three.js script that renders an interactive 3D robot."
          }
        ]
      }
    ],
    "generationConfig": {
      "maxOutputTokens": 8192
    }
  }'

cURL Code Example

#!/bin/bash

curl "https://api.cometapi.com/v1beta/models/gemini-3.5-flash-lite:generateContent" \
  -H "x-goog-api-key: $COMETAPI_KEY" \
  -H 'Content-Type: application/json' \
  -X POST \
  -d '{
    "contents": [
      {
        "parts": [
          {
            "text": "Write a three.js script that renders an interactive 3D robot."
          }
        ]
      }
    ],
    "generationConfig": {
      "maxOutputTokens": 8192
    }
  }'

Python Code Example

import os

from google import genai
from google.genai import types


client = genai.Client(
    api_key=os.environ["COMETAPI_KEY"],
    http_options={
        "api_version": "v1beta",
        "base_url": "https://api.cometapi.com",
        "retry_options": {"attempts": 1},
    },
)

response = client.models.generate_content(
    model="gemini-3.5-flash-lite",
    contents="Write a three.js script that renders an interactive 3D robot.",
    config=types.GenerateContentConfig(max_output_tokens=8192),
)

print(f"Response ID: {response.response_id}")
print(f"Model Version: {response.model_version}")
print(f"Finish Reason: {response.candidates[0].finish_reason}")
print(response.text)

JavaScript Code Example

import { GoogleGenAI } from "@google/genai";


const ai = new GoogleGenAI({
  apiKey: process.env.COMETAPI_KEY,
  httpOptions: {
    apiVersion: "v1beta",
    baseUrl: "https://api.cometapi.com",
    retryOptions: { attempts: 1 },
  },
});

const response = await ai.models.generateContent({
  model: "gemini-3.5-flash-lite",
  contents: "Write a three.js script that renders an interactive 3D robot.",
  config: {
    maxOutputTokens: 8192,
  },
});

console.log(`Response ID: ${response.responseId}`);
console.log(`Model Version: ${response.modelVersion}`);
console.log(`Finish Reason: ${response.candidates[0].finishReason}`);
console.log(response.text);

Uptime

Taux de succรจs des requรชtes sur les 30 derniers jours, reflรฉtant la fiabilitรฉ de chaque fournisseur de modรจles. CometAPI surveille tous les fournisseurs connectรฉs en temps rรฉel, 24h/24 et 7j/7.

RespondLIVE
1655msAvg. Response
UptimeLIVE
100.0%Avg. Uptime

Versions de Gemini 3.5 Flash Lite

La raison pour laquelle Gemini 3.5 Flash Lite dispose de plusieurs instantanรฉs peut inclure des facteurs potentiels tels que des variations de sortie aprรจs des mises ร  jour nรฉcessitant des instantanรฉs plus anciens pour la cohรฉrence, offrant aux dรฉveloppeurs une pรฉriode de transition pour l'adaptation et la migration, et diffรฉrents instantanรฉs correspondant ร  des points de terminaison globaux ou rรฉgionaux pour optimiser l'expรฉrience utilisateur. Pour les diffรฉrences dรฉtaillรฉes entre les versions, veuillez consulter la documentation officielle.

Version
gemini-3.5-flash-lite