모델
모델
wan2.6
wan2.6
veo3.1-lite
veo3.1-lite
Lightweight base variant of Veo3.1 with lower computing overhead, meets basic video generation demands for low-cost bulk preview scenarios.
doubao-seedance-2-0-fast
doubao-seedance-2-0-fast
Accelerated variant of Seedance 2.0. Balances generation speed and visual quality, reduces inference latency for high-concurrency testing and mass rapid content production.
doubao-seedance-2-0-mini
doubao-seedance-2-0-mini
Doubao Seedance 2.0 Mini is ByteDance’s lightweight text-to-video model under the Seed series. Optimized for fast inference and low GPU overhead, it generates smooth, style-consistent short videos from text prompts. Designed for API integration, test environments and low-cost batch video production.

GLM 5.1
TEST GLM-5.1 (released April 2026), purpose-built for long-horizon autonomous tasks. Unlike traditional models optimized for short interactions, GLM-5.1 excels at maintaining goal alignment, reducing strategy drift, and delivering production-grade results over extended periods — up to 8 hours of continuous autonomous work on a single complex task. It represents a major leap in agentic engineering, shifting evaluation from single-turn intelligence to real-world sustained execution.
gpt-5.6
gpt-5.6
Nano Banana 2 lite
Nano Banana 2 lite
Gemini 3.1 Flash Lite Image model은 이미지 생성 제품군에서 효율에 특화된 모델로, 초저지연과 비용 효율적인 이미지 생성 및 수정을 위해 설계되었습니다.
GPT-4.1 nano
GPT-4.1 nano
GPT-4.1 nano는 OpenAI에서 제공하는 인공지능 모델입니다. gpt-4.1-nano: 더 큰 컨텍스트 윈도우를 갖추었으며—최대 1 million 컨텍스트 토큰을 지원하고 향상된 긴 컨텍스트 이해를 통해 그 컨텍스트를 더 잘 활용할 수 있습니다. 지식 컷오프 시점은 2024년 6월로 업데이트되었습니다. 이 모델은 최대 1,047,576 토큰의 컨텍스트 길이를 지원합니다.

MiMo-V2.5
MiMo-V2.5 is Xiaomi's native full-modal model. It achieves professional-grade agent performance at about half the cost of inference, while outperforming MiMo-V2-Omni in multimodal perception in image and video understanding tasks.

GPT-5.4 pro
seedance-2-0
Happy Horse 1.1
Happy Horse 1.1
HappyHorse 1.1은 전문 콘텐츠 제작, 광고, 단편 영화, 소셜 미디어 제작 및 스토리텔링을 위해 설계된 멀티모달 비디오 생성 모델입니다. 이 모델은 독립적인 비디오 생성 평가에서 상위권을 기록해 큰 주목을 받은 HappyHorse 1.0의 역량을 더 강력한 장면 일관성과 향상된 시각적 충실도로 확장합니다.

GPT Image 2
OpenAI의 가장 강력한 이미지 생성 모델로, 다양한 언어에서 거의 완벽한 텍스트 렌더링과 최대 4K 해상도, 추론 기반 Thinking Mode를 제공합니다. 정확성, 속도, 그리고 브랜드에 부합하는 시각적 출력이 요구되는 프로덕션 워크플로에 맞춰 설계되었습니다.

DeepSeek V4 Flash
DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and high-throughput workloads, while maintaining strong reasoning and coding performance.
.jpeg&w=3840&q=75)
GPT 5.5
복잡한 코딩, 에이전트 기반 워크플로, 컴퓨터 활용, 데이터 분석 및 심층 연구를 위해 설계된 OpenAI의 가장 지능적이고 직관적인 플래그십 모델입니다. GPT-5.4와 같은 낮은 지연 시간으로 최첨단 수준의 지능을 제공하며, 더 높은 토큰 효율성으로 작업을 수행합니다. 까다로운 전문 및 엔터프라이즈 워크로드를 위한 최우선 선택 모델입니다.

GPT 5.5 Pro
OpenAI의 가장 역량이 뛰어난 모델로, 가장 어려운 작업과 장시간 실행되는 에이전트 기반 워크플로우를 위해 설계되었습니다. GPT-5.5 Pro는 복잡한 코딩, 컴퓨터 활용, 심층 연구, 데이터 분석, 과학적 추론에서 뛰어난 성능을 발휘하며, 더 높은 토큰 효율성과 함께 GPT-5.4 수준의 지연 시간으로 최첨단 수준의 지능을 제공합니다. 최고 수준의 정확성과 자율적 작업 실행을 요구하는 기업 및 전문 분야의 사용 사례에 이상적입니다.

Claude Sonnet 4.6
Claude Sonnet 4.6 is our most capable Sonnet model yet. It’s a full upgrade of the model’s skills across coding, computer use, long-context reasoning, agent planning, knowledge work, and design. Sonnet 4.6 also features a 1M token context window in beta.
Grok Imagine Video 1.5
Grok Imagine Video 1.5
xAI의 최신 이미지-투-비디오 모델은 네이티브 동기화 오디오 생성 기능을 갖추고 있으며 — 영상과 오디오를 한 번의 추론 패스로 함께 생성합니다. 480p/720p 출력과 최대 15초 길이의 클립을 지원하고, Image-to-Video Arena 리더보드에서 #1을 기록했습니다.

Claude Opus 4.7
Claude Opus 4.7 is a hybrid reasoning model designed specifically for frontier-level coding, AI agents, and complex multi-step professional work. Unlike lighter models (e.g., Sonnet or Haiku variants), Opus 4.7 prioritizes depth, consistency, and autonomy on the hardest tasks.
Happy Horse 1.0
Happy Horse 1.0
Happy Horse 1.0 — A high-quality audio-video generation model that supports text-to-video and image-to-video creation. It can generate synchronized visuals, audio, and lip movements, making it suitable for short films, advertising creatives, and product showcases.

Grok 4.3
Excels at agentic reasoning, knowledge work, and tool use.

GPT Image 2 ALL
GPT Image 2 is openai state-of-the-art image generation model for fast, high-quality image generation and editing. It supports flexible image sizes and high-fidelity image inputs.

Doubao-Seedance-2-0
Seedance 2.0은 시네마틱하고 멀티샷 내러티브 비디오 생성에 초점을 맞춘 ByteDance의 차세대 멀티모달 비디오 파운데이션 모델입니다. 단일 샷 텍스트-투-비디오 데모와 달리, Seedance 2.0은 레퍼런스 기반 제어(이미지, 짧은 클립, 오디오), 샷 간 캐릭터/스타일의 일관성, 그리고 네이티브 오디오/비디오 동기화를 강조하며 — AI 비디오가 전문적인 크리에이티브 및 프리비주얼라이제이션 워크플로우에서 유용하게 쓰이도록 하는 것을 목표로 합니다.

Sora 2 Pro
Sora 2 Pro는 동기화된 오디오가 포함된 동영상을 생성할 수 있는, 당사에서 가장 진보되고 강력한 미디어 생성 모델입니다. 자연어 또는 이미지로부터 정교하고 역동적인 동영상 클립을 생성할 수 있습니다.

Nano Banana 2
Core Capabilities Overview: Resolution: Up to 4K (4096×4096), on par with Pro. Reference Image Consistency: Up to 14 reference images (10 objects + 4 characters), maintaining style/character consistency. Extreme Aspect Ratios: New 1:4, 4:1, 1:8, 8:1 ratios added, suitable for long images, posters, and banners. Text Rendering: Advanced text generation, suitable for infographics and marketing poster layouts. Search Enhancement: Integrated Google Search + Image Search. Grounding: Built-in thinking process; complex prompts are reasoned before generation.

GPT-5.2 Pro
gpt-5.2-pro is the highest-capability, production-oriented member of OpenAI’s GPT-5.2 family, exposed through the Responses API for workloads that demand maximal fidelity, multi-step reasoning, extensive tool use and the largest context/throughput budgets OpenAI offers.

DeepSeek V4 Pro
DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding, and long-horizon agent workflows, with strong performance across knowledge, math, and software engineering benchmarks.

MiniMax-M2.7
MiniMax-M2.7 offers the same top-tier intelligence as the standard version—including recursive self-evolution and expert-level office productivity—but is designed for applications requiring sub-second latency and high-speed token generation. Leveraging an enhanced inference backbone architecture, its output speed is 66% faster than the standard model (reaching 100 tps). It is the preferred choice for interactive programming assistants, real-time agent loop execution, and high-throughput enterprise pipelines with stringent completion time requirements.

GPT-5.4 nano
GPT-5.4 nano is designed for tasks where speed and cost matter most like classification, data extraction, ranking, and sub-agents.

GPT-5.4 mini
GPT-5.4 mini brings the strengths of GPT-5.4 to a faster, more efficient model designed for high-volume workloads.
Gemini omni fast
Gemini omni fast
Omni is the new model that can create anything from any input — starting with video. With Omni, you can combine images, audio, video and text as input and generate high-quality videos grounded in Gemini's real-world knowledge. You can also easily edit your videos through conversation.