模型
模型
wan2.6
wan2.6
veo3.1-lite
veo3.1-lite
Lightweight base variant of Veo3.1 with lower computing overhead, meets basic video generation demands for low-cost bulk preview scenarios.
doubao-seedance-2-0-fast
doubao-seedance-2-0-fast
Accelerated variant of Seedance 2.0. Balances generation speed and visual quality, reduces inference latency for high-concurrency testing and mass rapid content production.
doubao-seedance-2-0-mini
doubao-seedance-2-0-mini
Doubao Seedance 2.0 Mini is ByteDance’s lightweight text-to-video model under the Seed series. Optimized for fast inference and low GPU overhead, it generates smooth, style-consistent short videos from text prompts. Designed for API integration, test environments and low-cost batch video production.

GLM 5.1
TEST GLM-5.1 (released April 2026), purpose-built for long-horizon autonomous tasks. Unlike traditional models optimized for short interactions, GLM-5.1 excels at maintaining goal alignment, reducing strategy drift, and delivering production-grade results over extended periods — up to 8 hours of continuous autonomous work on a single complex task. It represents a major leap in agentic engineering, shifting evaluation from single-turn intelligence to real-world sustained execution.
gpt-5.6
gpt-5.6
Nano Banana 2 lite
Nano Banana 2 lite
Gemini 3.1 Flash Lite Image model 是圖像生成系列中的效率專家,專為超低延遲與具成本效益的圖像生成與修改而設計。
GPT-4.1 nano
GPT-4.1 nano
GPT-4.1 nano 是由 OpenAI 提供的人工智慧模型。 gpt-4.1-nano: 具備更大的上下文視窗—支援最多 1 million 個上下文 token,並能透過改進的長上下文理解更好地利用該上下文。 知識截止時間更新為 2024 年 6 月。 此模型支援的最大上下文長度為 1,047,576 個 token。

MiMo-V2.5
MiMo-V2.5 is Xiaomi's native full-modal model. It achieves professional-grade agent performance at about half the cost of inference, while outperforming MiMo-V2-Omni in multimodal perception in image and video understanding tasks.

GPT-5.4 pro
seedance-2-0
Happy Horse 1.1
Happy Horse 1.1
HappyHorse 1.1 是一款多模態影片生成模型,專為專業內容製作、廣告、短片、社群媒體製作與故事創作而設計。它在 HappyHorse 1.0 的基礎上擴展了能力——HappyHorse 1.0 因在獨立的影片生成評測中名列前茅而廣受關注——並具備更強的場景連貫性與更佳的視覺保真度。

GPT Image 2
OpenAI 最強大的圖像生成模型,具備跨多語言近乎完美的文字渲染、最高可達 4K 解析度,以及由推理驅動的 Thinking Mode。專為對準確性、速度與符合品牌調性的視覺輸出有嚴格要求的生產級工作流程而打造。

DeepSeek V4 Flash
DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and high-throughput workloads, while maintaining strong reasoning and coding performance.
.jpeg&w=3840&q=75)
GPT 5.5
OpenAI 最智慧且最直觀的旗艦模型,專為複雜程式設計、代理式工作流程、電腦操作、資料分析與深入研究而設計。以與 GPT-5.4 相同的低延遲提供前沿級智慧,同時以更高的 Token 效率完成任務。是面向高要求的專業與企業級工作負載的首選模型。

GPT 5.5 Pro
OpenAI 最強大的模型,專為最艱鉅的任務與長時間運行的代理式工作流程打造。GPT-5.5 Pro 在複雜程式設計、電腦操作、深度研究、資料分析與科學推理方面表現卓越—以 GPT-5.4 的延遲並具備更高的 token 效率,提供前沿級智能。非常適合對準確性與自主任務執行提出最高標準要求的企業與專業應用場景。

Claude Sonnet 4.6
Claude Sonnet 4.6 is our most capable Sonnet model yet. It’s a full upgrade of the model’s skills across coding, computer use, long-context reasoning, agent planning, knowledge work, and design. Sonnet 4.6 also features a 1M token context window in beta.
Grok Imagine Video 1.5
Grok Imagine Video 1.5
xAI 最新的圖像轉影片模型具備原生同步音訊生成功能 — 影片與音訊可在單次推論中同時產生。支援 480p/720p 輸出,片段最長 15 秒,並在 Image-to-Video Arena 排行榜上排名 #1。

Claude Opus 4.7
Claude Opus 4.7 is a hybrid reasoning model designed specifically for frontier-level coding, AI agents, and complex multi-step professional work. Unlike lighter models (e.g., Sonnet or Haiku variants), Opus 4.7 prioritizes depth, consistency, and autonomy on the hardest tasks.
Happy Horse 1.0
Happy Horse 1.0
Happy Horse 1.0 — A high-quality audio-video generation model that supports text-to-video and image-to-video creation. It can generate synchronized visuals, audio, and lip movements, making it suitable for short films, advertising creatives, and product showcases.

Grok 4.3
Excels at agentic reasoning, knowledge work, and tool use.

GPT Image 2 ALL
GPT Image 2 is openai state-of-the-art image generation model for fast, high-quality image generation and editing. It supports flexible image sizes and high-fidelity image inputs.

Doubao-Seedance-2-0
Seedance 2.0 是 ByteDance 面向電影級、多鏡頭敘事影片生成的新一代多模態影片基礎模型。不同於單鏡頭的文字轉影片示範,Seedance 2.0 強調基於參考的控制(圖像、短片段、音訊)、角色與風格在跨鏡頭間的連貫與一致性,以及原生的聲畫同步——旨在讓 AI 影片切實服務於專業創作與前期預視工作流程。

Sora 2 Pro
Sora 2 Pro 是我們最先進且最強大的媒體生成模型,能生成帶有同步音訊的影片。它可以從自然語言或圖像創建細節豐富、動態的影片片段。

Nano Banana 2
Core Capabilities Overview: Resolution: Up to 4K (4096×4096), on par with Pro. Reference Image Consistency: Up to 14 reference images (10 objects + 4 characters), maintaining style/character consistency. Extreme Aspect Ratios: New 1:4, 4:1, 1:8, 8:1 ratios added, suitable for long images, posters, and banners. Text Rendering: Advanced text generation, suitable for infographics and marketing poster layouts. Search Enhancement: Integrated Google Search + Image Search. Grounding: Built-in thinking process; complex prompts are reasoned before generation.

GPT-5.2 Pro
gpt-5.2-pro is the highest-capability, production-oriented member of OpenAI’s GPT-5.2 family, exposed through the Responses API for workloads that demand maximal fidelity, multi-step reasoning, extensive tool use and the largest context/throughput budgets OpenAI offers.

DeepSeek V4 Pro
DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding, and long-horizon agent workflows, with strong performance across knowledge, math, and software engineering benchmarks.

MiniMax-M2.7
MiniMax-M2.7 offers the same top-tier intelligence as the standard version—including recursive self-evolution and expert-level office productivity—but is designed for applications requiring sub-second latency and high-speed token generation. Leveraging an enhanced inference backbone architecture, its output speed is 66% faster than the standard model (reaching 100 tps). It is the preferred choice for interactive programming assistants, real-time agent loop execution, and high-throughput enterprise pipelines with stringent completion time requirements.

GPT-5.4 nano
GPT-5.4 nano is designed for tasks where speed and cost matter most like classification, data extraction, ranking, and sub-agents.

GPT-5.4 mini
GPT-5.4 mini brings the strengths of GPT-5.4 to a faster, more efficient model designed for high-volume workloads.
Gemini omni fast
Gemini omni fast
Omni is the new model that can create anything from any input — starting with video. With Omni, you can combine images, audio, video and text as input and generate high-quality videos grounded in Gemini's real-world knowledge. You can also easily edit your videos through conversation.