Modelle
Modelle
wan2.6
wan2.6
veo3.1-lite
veo3.1-lite
Lightweight base variant of Veo3.1 with lower computing overhead, meets basic video generation demands for low-cost bulk preview scenarios.
doubao-seedance-2-0-fast
doubao-seedance-2-0-fast
Accelerated variant of Seedance 2.0. Balances generation speed and visual quality, reduces inference latency for high-concurrency testing and mass rapid content production.
doubao-seedance-2-0-mini
doubao-seedance-2-0-mini
Doubao Seedance 2.0 Mini is ByteDance’s lightweight text-to-video model under the Seed series. Optimized for fast inference and low GPU overhead, it generates smooth, style-consistent short videos from text prompts. Designed for API integration, test environments and low-cost batch video production.

GLM 5.1
TEST GLM-5.1 (released April 2026), purpose-built for long-horizon autonomous tasks. Unlike traditional models optimized for short interactions, GLM-5.1 excels at maintaining goal alignment, reducing strategy drift, and delivering production-grade results over extended periods — up to 8 hours of continuous autonomous work on a single complex task. It represents a major leap in agentic engineering, shifting evaluation from single-turn intelligence to real-world sustained execution.
gpt-5.6
gpt-5.6
Nano Banana 2 lite
Nano Banana 2 lite
Das Modell Gemini 3.1 Flash Lite Image ist ein Effizienzspezialist in der Modellfamilie für Bildgenerierung und wurde für ultraniedrige Latenz sowie kosteneffiziente Bildgenerierung und -bearbeitung entwickelt.
GPT-4.1 nano
GPT-4.1 nano
GPT-4.1 nano ist ein von OpenAI bereitgestelltes KI-Modell. gpt-4.1-nano: Verfügt über ein größeres Kontextfenster—unterstützt bis zu 1 Million Kontext-Token und kann diesen Kontext dank eines verbesserten Verständnisses langer Kontexte besser nutzen. Verfügt über einen aktualisierten Wissensstand bis Juni 2024. Dieses Modell unterstützt eine maximale Kontextlänge von 1,047,576 Token.

MiMo-V2.5
MiMo-V2.5 is Xiaomi's native full-modal model. It achieves professional-grade agent performance at about half the cost of inference, while outperforming MiMo-V2-Omni in multimodal perception in image and video understanding tasks.

GPT-5.4 pro
seedance-2-0
Happy Horse 1.1
Happy Horse 1.1
HappyHorse 1.1 ist ein multimodales Modell zur Videogenerierung, das für professionelle Inhaltserstellung, Werbung, Kurzfilme, Social-Media-Produktion und Storytelling konzipiert ist. Es erweitert die Fähigkeiten von HappyHorse 1.0 — das nach hohen Platzierungen in unabhängigen Evaluierungen zur Videogenerierung große Aufmerksamkeit erlangte — um eine stärkere Szenenkohärenz und eine verbesserte Bildtreue.

GPT Image 2
OpenAIs leistungsfähigstes Bildgenerierungsmodell mit nahezu perfekter Textdarstellung in mehreren Sprachen, bis zu 4K-Auflösung und Reasoning-gestütztem Thinking Mode. Entwickelt für Produktions-Workflows, die Genauigkeit, Geschwindigkeit und markenkonforme visuelle Ergebnisse verlangen.

DeepSeek V4 Flash
DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and high-throughput workloads, while maintaining strong reasoning and coding performance.
.jpeg&w=3840&q=75)
GPT 5.5
Das intelligenteste und intuitivste Flaggschiffmodell von OpenAI, entwickelt für komplexes Programmieren, agentische Workflows, Computer-Nutzung, Datenanalyse und tiefgehende Forschung. Liefert Intelligenz auf Frontier-Niveau bei gleicher niedriger Latenz wie GPT-5.4, während es Aufgaben mit höherer Token-Effizienz erledigt. Das bevorzugte Modell für anspruchsvolle professionelle und unternehmensweite Workloads.

GPT 5.5 Pro
OpenAIs leistungsfähigstes Modell, entwickelt für die anspruchsvollsten Aufgaben und langandauernde agentische Workflows. GPT-5.5 Pro überzeugt bei komplexem Programmieren, der Computernutzung, tiefgehender Recherche, der Datenanalyse und wissenschaftlichem Schlussfolgern — und liefert Intelligenz auf Spitzenniveau bei GPT-5.4-Latenz mit höherer Token-Effizienz. Ideal für Unternehmens- und professionelle Anwendungsfälle, die höchste Präzision und autonome Aufgabenausführung verlangen.

Claude Sonnet 4.6
Claude Sonnet 4.6 is our most capable Sonnet model yet. It’s a full upgrade of the model’s skills across coding, computer use, long-context reasoning, agent planning, knowledge work, and design. Sonnet 4.6 also features a 1M token context window in beta.
Grok Imagine Video 1.5
Grok Imagine Video 1.5
Das neueste Bild-zu-Video-Modell von xAI bietet eine native, synchronisierte Audiogenerierung — Video und Ton werden in einem einzigen Inferenzdurchlauf erzeugt. Unterstützt 480p/720p-Ausgabe, Clips von bis zu 15 Sekunden und rangiert auf Platz 1 in der Rangliste der Image-to-Video Arena.

Claude Opus 4.7
Claude Opus 4.7 is a hybrid reasoning model designed specifically for frontier-level coding, AI agents, and complex multi-step professional work. Unlike lighter models (e.g., Sonnet or Haiku variants), Opus 4.7 prioritizes depth, consistency, and autonomy on the hardest tasks.
Happy Horse 1.0
Happy Horse 1.0
Happy Horse 1.0 — A high-quality audio-video generation model that supports text-to-video and image-to-video creation. It can generate synchronized visuals, audio, and lip movements, making it suitable for short films, advertising creatives, and product showcases.

Grok 4.3
Excels at agentic reasoning, knowledge work, and tool use.

GPT Image 2 ALL
GPT Image 2 is openai state-of-the-art image generation model for fast, high-quality image generation and editing. It supports flexible image sizes and high-fidelity image inputs.

Doubao-Seedance-2-0
Seedance 2.0 ist das multimodale Video-Grundlagenmodell der nächsten Generation von ByteDance, das auf filmische, narrative Videogenerierung mit mehreren Einstellungen ausgerichtet ist. Im Gegensatz zu Text-zu-Video-Demos mit nur einer Einstellung legt Seedance 2.0 den Schwerpunkt auf referenzbasierte Steuerung (Bilder, kurze Clips, Audio), durchgängige Konsistenz von Charakteren und Stil über mehrere Einstellungen hinweg sowie native Audio-/Video-Synchronisierung — mit dem Ziel, KI-Video für professionelle Kreativ- und Previsualisierungs-Workflows nutzbar zu machen.

Sora 2 Pro
Sora 2 Pro ist unser fortschrittlichstes und leistungsstärkstes Modell zur Mediengenerierung, das Videos mit synchronisiertem Audio generieren kann. Es kann aus natürlicher Sprache oder Bildern detaillierte, dynamische Videoclips generieren.

Nano Banana 2
Core Capabilities Overview: Resolution: Up to 4K (4096×4096), on par with Pro. Reference Image Consistency: Up to 14 reference images (10 objects + 4 characters), maintaining style/character consistency. Extreme Aspect Ratios: New 1:4, 4:1, 1:8, 8:1 ratios added, suitable for long images, posters, and banners. Text Rendering: Advanced text generation, suitable for infographics and marketing poster layouts. Search Enhancement: Integrated Google Search + Image Search. Grounding: Built-in thinking process; complex prompts are reasoned before generation.

GPT-5.2 Pro
gpt-5.2-pro is the highest-capability, production-oriented member of OpenAI’s GPT-5.2 family, exposed through the Responses API for workloads that demand maximal fidelity, multi-step reasoning, extensive tool use and the largest context/throughput budgets OpenAI offers.

DeepSeek V4 Pro
DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding, and long-horizon agent workflows, with strong performance across knowledge, math, and software engineering benchmarks.

MiniMax-M2.7
MiniMax-M2.7 offers the same top-tier intelligence as the standard version—including recursive self-evolution and expert-level office productivity—but is designed for applications requiring sub-second latency and high-speed token generation. Leveraging an enhanced inference backbone architecture, its output speed is 66% faster than the standard model (reaching 100 tps). It is the preferred choice for interactive programming assistants, real-time agent loop execution, and high-throughput enterprise pipelines with stringent completion time requirements.

GPT-5.4 nano
GPT-5.4 nano is designed for tasks where speed and cost matter most like classification, data extraction, ranking, and sub-agents.

GPT-5.4 mini
GPT-5.4 mini brings the strengths of GPT-5.4 to a faster, more efficient model designed for high-volume workloads.
Gemini omni fast
Gemini omni fast
Omni is the new model that can create anything from any input — starting with video. With Omni, you can combine images, audio, video and text as input and generate high-quality videos grounded in Gemini's real-world knowledge. You can also easily edit your videos through conversation.