Unlock exclusive introductory pricing for the newly launched Gemini 3.5 Flash.
O

GPT-4o mini Realtime Preview

입력:$60/M
출력:$60/M
출시일:Oct 1, 2025

GPT-4o mini Realtime Preview는 대화형 음성 및 시각적 경험을 위한 실시간 멀티모달 모델입니다. 스트리밍 입력과 출력으로 음성, 텍스트, 이미지를 처리하며, 구체적인 동작 수행을 위한 툴/함수 호출도 지원합니다. 대표적인 활용 사례로는 음성 비서, 실시간 통화 처리, 실시간 자막 생성, 카메라 또는 화면 콘텐츠를 통한 시각적 질의응답이 있습니다. 기술적 하이라이트로는 양방향 오디오, 시각 이해, 스트리밍 응답, 함수 기반의 구조화된 출력이 포함됩니다.

인기
상업적 사용

Technical Specifications of gpt-4o-mini-realtime-preview

SpecificationDetails
Model IDgpt-4o-mini-realtime-preview
ProviderOpenAI via CometAPI
ModalitiesText, audio, image
Input typesStreaming audio, text messages, image inputs
Output typesStreaming text, synthesized/streamed audio, structured function calls
Core strengthsLow-latency interaction, multimodal understanding, real-time conversation, tool use
Best forVoice assistants, live support calls, captioning, visual Q&A, interactive agents
Function callingSupported
StreamingSupported
Realtime sessionsSupported
Typical interaction patternContinuous bidirectional session with incremental input and output

What is gpt-4o-mini-realtime-preview?

gpt-4o-mini-realtime-preview is a real-time multimodal model designed for fast, interactive experiences where users speak, type, or share visual input and expect immediate responses. It is well suited for applications that need live back-and-forth communication rather than standard single-turn request/response workflows.

The model can process speech, text, and images within the same experience, making it useful for assistants that listen to a caller, inspect on-screen or camera content, and respond in natural language or audio. Because it supports streaming input and output, developers can build systems that feel responsive during ongoing interactions instead of waiting for a full completion.

It also supports tool or function calling, which allows the model to trigger structured actions such as looking up data, calling backend services, or executing workflow steps. This makes gpt-4o-mini-realtime-preview a strong choice for grounded, action-oriented agents in customer support, operations, productivity, and multimodal assistant scenarios.

Main features of gpt-4o-mini-realtime-preview

  • Real-time multimodal interaction: Accepts and responds across speech, text, and images for fluid live experiences.
  • Bidirectional audio: Supports conversational voice interfaces where audio can be streamed in and responses can be streamed back out.
  • Streaming responses: Delivers partial outputs incrementally, reducing perceived latency and improving responsiveness.
  • Vision understanding: Interprets visual inputs such as camera frames, screenshots, or other images during a live session.
  • Function and tool calling: Produces structured calls that let your application connect the model to business logic, databases, or external tools.
  • Interactive agent behavior: Works well for assistants that must maintain turn-by-turn context during active sessions.
  • Live call handling: Useful for phone or web-call scenarios involving fast speech understanding and immediate replies.
  • Real-time captioning and transcription workflows: Can support experiences that convert ongoing speech into usable text in near real time.
  • Structured outputs for actions: Helps applications turn conversational intent into reliable machine-readable instructions.
  • Low-latency user experiences: Optimized for scenarios where responsiveness matters, such as support, coaching, monitoring, and guided workflows.

How to access and integrate gpt-4o-mini-realtime-preview

Step 1: Sign Up for API Key

First, create an account on CometAPI and generate your API key from the dashboard. This key is required to authenticate every request. Store it securely and avoid exposing it in client-side code or public repositories.

Step 2: Connect to gpt-4o-mini-realtime-preview API

The Realtime API uses WebSocket connections. Connect to CometAPI's WebSocket endpoint:

const ws = new WebSocket(
  "wss://api.cometapi.com/v1/realtime?model=gpt-4o-mini-realtime-preview",
  {
    headers: {
      "Authorization": "Bearer " + process.env.COMETAPI_API_KEY,
      "OpenAI-Beta": "realtime=v1"
    }
  }
);

ws.on("open", () => {
  ws.send(JSON.stringify({
    type: "session.update",
    session: {
      modalities: ["text", "audio"],
      instructions: "You are a helpful assistant."
    }
  }));
});

ws.on("message", (data) => {
  console.log(JSON.parse(data));
});

Step 3: Retrieve and Verify Results

The Realtime API streams responses through the WebSocket connection as server-sent events. Listen for response.audio.delta events for audio output and response.text.delta for text. Verify the session is established and responses are streaming correctly.

GPT-4o mini Realtime Preview 가격

[모델명]의 경쟁력 있는 가격을 살펴보세요. 다양한 예산과 사용 요구에 맞게 설계되었습니다. 유연한 요금제로 사용한 만큼만 지불하므로 요구사항이 증가함에 따라 쉽게 확장할 수 있습니다. [모델명]이 비용을 관리 가능한 수준으로 유지하면서 프로젝트를 어떻게 향상시킬 수 있는지 알아보세요.

ModelComet Price (USD / M Tokens)Official Price (USD / M Tokens)Discount
gpt-4o-mini-realtime-preview
입력:$0.48/M
출력:$1.92/M
입력:$0.6/M
출력:$2.4/M
-20%
gpt-4o-mini-realtime-preview-2024-12-17
입력:$0.48/M
출력:$1.92/M
입력:$0.6/M
출력:$2.4/M
-20%

GPT-4o mini Realtime Preview의 샘플 코드 및 API

[모델 이름]의 포괄적인 샘플 코드와 API 리소스에 액세스하여 통합 프로세스를 간소화하세요. 자세한 문서는 단계별 가이드를 제공하여 프로젝트에서 [모델 이름]의 모든 잠재력을 활용할 수 있도록 돕습니다.

GPT-4o mini Realtime Preview의 버전

GPT-4o mini Realtime Preview에 여러 스냅샷이 존재하는 이유는 업데이트 후 출력 변동으로 인해 일관성을 유지하기 위해 이전 스냅샷을 보관하거나, 개발자에게 적응 및 마이그레이션을 위한 전환 기간을 제공하거나, 글로벌 또는 지역별 엔드포인트에 따라 다양한 스냅샷을 제공하여 사용자 경험을 최적화하기 위한 것 등이 포함될 수 있습니다. 버전 간 상세한 차이점은 공식 문서를 참고해 주시기 바랍니다.

Version
gpt-4o-mini-realtime-preview
gpt-4o-mini-realtime-preview-2024-12-17