# Modellix - [Set Up AI Agents on Modellix](https://docs.modellix.ai/get-started/index.md): Set up an AI agent with one Modellix API key for Media Models, the LLM gateway, and optional Web Tools. Use async polling for media; call LLM and Tools synchronously. - [Modellix Platform Overview](https://docs.modellix.ai/get-started/overview.md): Learn how Modellix provides unified API access to Media Models and LLMs, plus optional Web Tools, with one API key, public pricing, and browser playgrounds. - [Modellix API Pricing for Media, LLM, and Tools](https://docs.modellix.ai/get-started/pricing.md): Understand Modellix pay-as-you-go pricing: media generation by image, second, or character; LLM by token; optional Web Tools by request or URL. - [Fund Your Modellix Account](https://docs.modellix.ai/get-started/top-up.md): Learn how to fund your Modellix account with supported payment methods, pay-as-you-go top-ups, and available discounts for your team. - [Modellix Rate Limits and Team Entitlements](https://docs.modellix.ai/get-started/entitlements.md): Understand Modellix concurrency limits, rate limits (RPM), and team entitlements that scale with your single top-up funding tier amount. - [Choose the Right Modellix Image, Video, or Speech Model](https://docs.modellix.ai/get-started/model-select.md): Choose the right Modellix image, video, or speech model by output type, quality, speed, cost, and features such as audio, references, and editing. - [Modellix AI Providers and Service URLs](https://docs.modellix.ai/get-started/model-providers.md): Browse every upstream provider Modellix routes through for Media Models, LLM, and Web Tools, including official service URLs and links to the matching API docs. - [Use the Modellix REST API for Media Generation](https://docs.modellix.ai/ways-to-use/api.md): Use the Modellix REST API to authenticate, upload media inputs, submit an asynchronous image, video, or audio task, poll its status, and retrieve output assets. - [Install the Modellix Agent Skill](https://docs.modellix.ai/ways-to-use/skill.md): Install the Modellix Agent Skill so coding agents in Cursor, Claude Code, and other hosts can discover models, inspect request schemas, and generate images, videos, and speech. - [Modellix Plugin for Coding Agents](https://docs.modellix.ai/ways-to-use/plugin.md): Install the official Modellix Plugin in Claude Code, Codex, Cursor, and other agent hosts to generate images, videos, and speech via CLI or REST API. - [Use Modellix Agent Canvas in Coding Agents](https://docs.modellix.ai/ways-to-use/agent-canvas.md): Install the Modellix Agent Canvas local stdio MCP plugin in Codex, Cursor, Claude Code, or OpenCode for visual image work, HTML drafts, and presentations. - [Use Modellix with the DeepSeek Harness Plugin](https://docs.modellix.ai/ways-to-use/deepseek-harness.md): Install the dsh-modellix plugin in DeepSeek Harness to use Modellix LLM models, media generation, and Web Search and Web Fetch with one API key. - [Modellix CLI for Image, Video, and Audio Generation](https://docs.modellix.ai/ways-to-use/cli.md): Use modellix-cli to authenticate, discover models, inspect request schemas, submit async image, video, and audio tasks, wait for results, and download assets from the terminal or CI. - [Docs](https://docs.modellix.ai/ways-to-use/docs-search-mcp.md): Connect AI clients like Cursor and Claude Desktop to the Docs MCP to search the official documentation without leaving your editor. - [Modellix Media Generation](https://docs.modellix.ai/ways-to-use/media-generation-mcp.md): Modellix Media Generation is an MCP server for image, video, and speech generation. It is in development. - [Generate Images, Videos, and Audio in the Playground](https://docs.modellix.ai/ways-to-use/playground.md): Generate images, videos, and speech audio in the Modellix Playground with no code required. Discover models, tune prompts, and preview outputs in your browser. - [Validate API Key](https://docs.modellix.ai/api/validate-api-key.md): Checks whether the API Key provided in the Authorization header is valid. Invalid, missing, or malformed credentials return a successful response with is_valid set to false. - [Get Team Balance](https://docs.modellix.ai/api/get-team-balance.md): Returns the current available balance for the team associated with the provided API Key. The balance is returned in USD with 4 decimal places. - [HappyHorse 1.1 T2V](https://docs.modellix.ai/alibaba/happyhorse-1-1-t2v.md): [Core Function] HappyHorse 1.1 T2V is Alibaba's latest streamlined text-to-video model. [Strengths] It generates 720P/1080P video with native audio support, 3-15 second duration, and an expanded set of aspect ratios including 4:5, 5:4, 9:21, and 21:9. [Best For] Highly recommended for: fast HappyHor… - [HappyHorse 1.1 I2V](https://docs.modellix.ai/alibaba/happyhorse-1-1-i2v.md): [Core Function] HappyHorse 1.1 I2V is Alibaba's latest streamlined first-frame image-to-video model. [Strengths] It turns a single image into high-quality 720P/1080P video with native audio support and 3-15 second duration; output aspect ratio follows the first frame image. [Best For] Highly recomme… - [HappyHorse 1.1 R2V](https://docs.modellix.ai/alibaba/happyhorse-1-1-r2v.md): [Core Function] HappyHorse 1.1 R2V is Alibaba's latest reference-image-to-video model. [Strengths] It uses 1-9 reference images to preserve subject or character appearance while generating new video actions, supports 720P/1080P output, 3-15 second duration, and expanded aspect ratios including 4:5,… - [HappyHorse 1.0 T2V](https://docs.modellix.ai/alibaba/happyhorse-1-0-t2v.md): [Core Function] HappyHorse 1.0 T2V is a breakout, highly optimized text-to-video model. [Strengths] It provides streamlined, fast, and high-quality video generation (up to 15s at 1080p) with native audio support, acting as a highly efficient alternative to Wan 2.7. [Best For] Highly recommended for:… - [HappyHorse 1.0 I2V](https://docs.modellix.ai/alibaba/happyhorse-1-0-i2v.md): [Core Function] HappyHorse 1.0 I2V is a streamlined image-to-video model. [Strengths] It generates high-quality 720P/1080P video (3-15s) from an image efficiently, with native audio support. [Best For] Highly recommended for: rapid image animation and robust character motion. [Limitations] Does not… - [HappyHorse 1.0 R2V](https://docs.modellix.ai/alibaba/happyhorse-1-0-r2v.md): [Core Function] HappyHorse 1.0 R2V is a reference-to-video model. [Strengths] It excels at maintaining character consistency using up to 9 reference images while generating new video actions based on a prompt. [Best For] Highly recommended for: character-consistent storytelling and generating multip… - [HappyHorse 1.0 Video Edit](https://docs.modellix.ai/alibaba/happyhorse-1-0-video-edit.md): [Core Function] HappyHorse 1.0 Video Edit is a streamlined video editing model. [Strengths] It provides high-quality video editing capabilities (with or without reference images) within the highly optimized HappyHorse architecture. [Best For] Highly recommended for: fast, high-quality video modifica… - [Qwen Image 3.0 Pro](https://docs.modellix.ai/alibaba/qwen-image-3-0-pro.md): [Core Function] Qwen Image 3.0 Pro is Alibaba's latest text-to-image model with strong prompt following and photorealism. [Strengths] It supports free-form output size (width*height), optional negative prompts, intelligent prompt rewrite, batch generation of 1-6 images, and long structured prompts f… - [Qwen Image 3.0 Pro Edit](https://docs.modellix.ai/alibaba/qwen-image-3-0-pro-edit.md): [Core Function] Qwen Image 3.0 Pro Edit is an image-to-image editing model for instruction-based edits and multi-image fusion. [Strengths] It accepts 1-3 reference images plus an edit instruction, optional negative prompts, free-form output size (width*height), intelligent prompt rewrite, 1-6 output… - [Qwen Image 3.0](https://docs.modellix.ai/alibaba/qwen-image-3-0.md): [Core Function] Qwen Image 3.0 is Alibaba's standard text-to-image model balancing quality and speed. [Strengths] It supports free-form output size (width*height), optional negative prompts, intelligent prompt rewrite (direct/agent modes), batch generation of 1-6 images, and long structured prompts.… - [Qwen Image 3.0 Edit](https://docs.modellix.ai/alibaba/qwen-image-3-0-edit.md): [Core Function] Qwen Image 3.0 Edit is the standard image-to-image editing model for instruction-based edits and multi-image fusion. [Strengths] It accepts 1-3 reference images plus an edit instruction, optional negative prompts, free-form output size (width*height), intelligent prompt rewrite (dire… - [Wan 3.0 T2V](https://docs.modellix.ai/alibaba/wan3-0-t2v.md): [Core Function] Wan 3.0 T2V is Alibaba Wan 3.0 text-to-video generation with optional document or webpage reference. [Strengths] It generates up to 30-second video at 480P/720P/1080P with controllable aspect ratio and optional output audio, and can ground generation on a file or public webpage. [Bes… - [Wan 3.0 I2V](https://docs.modellix.ai/alibaba/wan3-0-i2v.md): [Core Function] Wan 3.0 I2V is Alibaba Wan 3.0 image-to-video generation supporting first-frame, first-last-frame, and reference-image modes. [Strengths] It can strictly lock the first and last frames or fuse up to 10 reference images with optional reference audio for multimodal guidance. [Best For]… - [Wan 3.0 V2V](https://docs.modellix.ai/alibaba/wan3-0-v2v.md): [Core Function] Wan 3.0 V2V is Alibaba Wan 3.0 reference-video generation that builds new video from one or more input videos. [Strengths] It supports up to 5 reference videos with optional reference images and audio for multimodal composition and prompt-referenced subjects. [Best For] Highly recomm… - [Wan 2.7 T2V](https://docs.modellix.ai/alibaba/wan-2-7-t2v.md): [Core Function] Wan 2.7 T2V is Alibaba's flagship text-to-video generation model. [Strengths] It generates high-fidelity video directly from text with support for custom aspect ratios, audio generation, and intricate semantic adherence. [Best For] Highly recommended for: high-quality commercial vide… - [Wan 2.7 I2V](https://docs.modellix.ai/alibaba/wan-2-7-i2v.md): [Core Function] Wan 2.7 I2V is Alibaba's flagship multimodal image-to-video model. [Strengths] It supports multimodal input (text, image, audio, video) for first-frame, start-and-end-frame (FL2V), and video continuation tasks. [Best For] Highly recommended for: complex image animation, cinematic tra… - [Wan 2.7 R2V](https://docs.modellix.ai/alibaba/wan-2-7-r2v.md): [Core Function] Wan 2.7 Reference-to-Video is a highly capable character/entity reference video model. [Strengths] It natively supports entity reference, voice customization, and playbook-based video generation from a single storyboard. [Best For] Highly recommended for: creating consistent video se… - [Wan 2.7 Videoedit](https://docs.modellix.ai/alibaba/wan-2-7-videoedit.md): [Core Function] Wan 2.7 Video Editing is an instruction-based video modification model. [Strengths] It supports complex video editing tasks like content replacement using reference images, and replicating actions, effects, and camera movements. [Best For] Highly recommended for: modifying existing v… - [Wan 2.7 Image Pro](https://docs.modellix.ai/alibaba/wan-2-7-image-pro.md): [Core Function] Wan 2.7 Image Pro is Alibaba's flagship reasoning-enhanced image generation model. [Strengths] It features built-in chain-of-thought reasoning (Thinking Mode), exceptional prompt accuracy, native 12-language text rendering, and generates ultra-high-resolution 4K images. [Best For] Hi… - [Wan 2.7 Image Pro Edit](https://docs.modellix.ai/alibaba/wan-2-7-image-pro-edit.md): [Core Function] Wan 2.7 Image Pro Edit is Alibaba's flagship reasoning-enhanced image editing model. [Strengths] It supports interactive editing, character-consistent multi-image generation, and complex multi-reference modifications with deep reasoning. [Best For] Highly recommended for: professiona… - [Wan 2.7 Image](https://docs.modellix.ai/alibaba/wan-2-7-image.md): [Core Function] Wan 2.7 Image is a fast, reasoning-enhanced image generation model. [Strengths] It includes the chain-of-thought reasoning and text rendering of the Pro version, but is optimized for speed, supporting up to 2K resolution. [Best For] Highly recommended for: fast iterations, conceptual… - [Wan 2.7 Image Edit](https://docs.modellix.ai/alibaba/wan-2-7-image-edit.md): [Core Function] Wan 2.7 Image Edit is a fast, reasoning-enhanced image editing model. [Strengths] Provides the robust editing capabilities of the Wan 2.7 architecture with faster turnaround times. [Best For] Highly recommended for: standard image modifications and style transfers. [Limitations] Do N… - [Z Image Turbo](https://docs.modellix.ai/alibaba/z-image-turbo.md): [Core Function] Z-Image Turbo is an older generation text-to-image model. [Strengths] Historically provided faster generation times and lower latency compared to its standard counterparts. Maintained for backward compatibility. [Best For] Existing legacy integrations that require the specific speed/… - [CosyVoice V3 Plus](https://docs.modellix.ai/alibaba/cosyvoice-v3-plus.md): [Core Function] CosyVoice v3 Plus is Alibaba's high-quality text-to-speech model. [Strengths] System voices (e.g. longanyang, longanhuan), SSML and LaTeX input, hot_fix pronunciation correction, AIGC watermark, and output in mp3, pcm, wav, or opus. System-voice instruction must use the fixed Chinese… - [CosyVoice V3 Flash](https://docs.modellix.ai/alibaba/cosyvoice-v3-flash.md): [Core Function] CosyVoice v3 Flash is Alibaba's low-latency text-to-speech model. [Strengths] Rich system voice catalog, fixed-format instruction on Instruct-capable system voices, SSML, hot_fix, AIGC watermark, Markdown filter (cloned voices only), and multiple audio formats with faster turnaround… - [CosyVoice Design](https://docs.modellix.ai/alibaba/cosyvoice-design.md): [Core Function] CosyVoice Design creates a temporary voice from a natural-language voice_prompt and synthesizes speech in one async request; only the final audio is returned. [Strengths] No voice enrollment management; language_hint (zh/en) applies to both design and synthesis; same prosody and form… - [CosyVoice Clone](https://docs.modellix.ai/alibaba/cosyvoice-clone.md): [Core Function] CosyVoice Clone clones a speaker from a public reference audio URL and synthesizes new speech in one async request; only the final audio is returned. [Strengths] No voice enrollment management; language_hint applies to both cloning and synthesis; SSML, hot_fix, and prosody controls.… - [Qwen Audio 3.0 TTS Plus](https://docs.modellix.ai/alibaba/qwen-audio-3-0-tts-plus.md): [Core Function] Qwen-Audio 3.0 TTS Plus is Alibaba's high-quality Qwen-Audio text-to-speech model on the same SpeechSynthesizer endpoint family as CosyVoice. [Strengths] Natural speech synthesis with voice, format, sample-rate, prosody, SSML, instruction, language_hint, and AIGC watermark controls;… - [Qwen Audio 3.0 TTS Flash](https://docs.modellix.ai/alibaba/qwen-audio-3-0-tts-flash.md): [Core Function] Qwen-Audio 3.0 TTS Flash is Alibaba's low-latency Qwen-Audio text-to-speech model on the same SpeechSynthesizer endpoint family as CosyVoice. [Strengths] Fast synthesis with voice, format, sample-rate, prosody, SSML, instruction, language_hint, and AIGC watermark controls; system voi… - [Fun ASR](https://docs.modellix.ai/alibaba/fun-asr.md): [Core Function] Fun-ASR transcribes a single public audio file asynchronously. [Strengths] Hot-word vocabulary, optional speaker diarization, channel selection, and language hints. [Best For] Batch transcription of recordings up to 12 hours. [Limitations] Do NOT send more than one file per request.… - [Fun ASR MTL](https://docs.modellix.ai/alibaba/fun-asr-mtl.md): [Core Function] Fun-ASR MTL is the multi-language variant for async recorded speech recognition. [Strengths] Same parameters as fun-asr with multi-language tuning. [Best For] Mixed-language or international audio archives. [Limitations] Do NOT send more than one file per request. [Routing] Choose fu… - [Seedance 2.5 T2V](https://docs.modellix.ai/bytedance/seedance-2-5-t2v.md): [Core Function] Seedance 2.5 T2V is ByteDance Dreamina Seedance 2.5 text-to-video generation. [Strengths] It generates longer clips up to 30 seconds at 480p/720p/1080p with optional mp4 or mov output and native audio generation. [Best For] Highly recommended for: longer-form text-to-video storytelli… - [Seedance 2.5 I2V](https://docs.modellix.ai/bytedance/seedance-2-5-i2v.md): [Core Function] Seedance 2.5 I2V is ByteDance Dreamina Seedance 2.5 image-to-video generation supporting first-frame, first-and-last-frame, and reference-image modes. [Strengths] It supports up to 30 reference images, optional reference audio, 480p/720p/1080p, 4-30 second duration, and mp4 or mov ou… - [Seedance 2.5 V2V](https://docs.modellix.ai/bytedance/seedance-2-5-v2v.md): [Core Function] Seedance 2.5 V2V is ByteDance Dreamina Seedance 2.5 video-to-video generation covering multimodal reference, video editing, and video extension. [Strengths] It accepts up to 10 reference videos, 30 reference images, and 10 audio clips, with 480p/720p/1080p, 4-30 second output, and mp… - [Seedance 2.0 T2V](https://docs.modellix.ai/bytedance/seedance-2-0-t2v.md): [Core Function] Seedance 2.0 T2V is ByteDance Dreamina Seedance 2.0 text-to-video. [Strengths] Supports 480p/720p/1080p/4k, 24 fps, 4-15s MP4 output. Text-only input — do not pass images, video, or audio. [Routing] Use for high-fidelity text-to-video when quality or 4k output is requested. - [Seedance 2.0 I2V](https://docs.modellix.ai/bytedance/seedance-2-0-i2v.md): [Core Function] Seedance 2.0 I2V is ByteDance's flagship unified multimodal video generation model. [Strengths] It supports complex mixed references (multiple images, audio clips) and generates up to 15s of multi-shot audio-video output with dual-channel audio. [Best For] Highly recommended for: hig… - [Seedance 2.0 V2V](https://docs.modellix.ai/bytedance/seedance-2-0-v2v.md): [Core Function] Seedance 2.0 V2V is ByteDance's flagship multimodal video-to-video model. [Strengths] It allows powerful editing and stylization of input videos by supporting mixed references (text, images, video, and audio) and producing multi-shot 15s outputs. [Best For] Highly recommended for: co… - [Seedance 2.0 Fast T2V](https://docs.modellix.ai/bytedance/seedance-2-0-fast-t2v.md): [Core Function] Seedance 2.0 Fast T2V is the faster text-to-video variant. [Strengths] Supports 480p/720p, 24 fps, 4-15s MP4 output. [Routing] Use when speed is preferred over maximum resolution. - [Seedance 2.0 Fast I2V](https://docs.modellix.ai/bytedance/seedance-2-0-fast-i2v.md): [Core Function] Seedance 2.0 Fast I2V is a high-speed multimodal video generation model. [Strengths] Fast generation with the multimodal and multi-shot capabilities of the Seedance 2.0 architecture. [Best For] Highly recommended for: rapid prototyping and quick multi-shot video creation. [Limitation… - [Seedance 2.0 Fast V2V](https://docs.modellix.ai/bytedance/seedance-2-0-fast-v2v.md): [Core Function] Seedance 2.0 Fast V2V is a high-speed video-to-video generation model. [Strengths] It offers rapid video transformation capabilities based on the 2.0 architecture. [Best For] Highly recommended for: quick video restyling and fast iterations. [Limitations] Do NOT use this model when a… - [Seedance 2.0 Mini T2V](https://docs.modellix.ai/bytedance/seedance-2-0-mini-t2v.md): [Core Function] Seedance 2.0 Mini T2V is the lightweight text-to-video variant. [Strengths] Supports 480p/720p, 24 fps, 4-15s MP4 output. [Routing] Use for cost-efficient text-to-video. - [Seedance 2.0 Mini I2V](https://docs.modellix.ai/bytedance/seedance-2-0-mini-i2v.md): [Core Function] Seedance 2.0 Mini I2V is the lightweight multimodal image-to-video variant. [Strengths] Supports 480p/720p, 24 fps, 4-15s MP4 output, with optional first/last frame, reference images, and audio references. [Routing] Use for cost-efficient image-to-video. - [Seedance 2.0 Mini V2V](https://docs.modellix.ai/bytedance/seedance-2-0-mini-v2v.md): [Core Function] Seedance 2.0 Mini V2V is the lightweight video-to-video variant. [Strengths] Supports 480p/720p, 24 fps, 4-15s MP4 output, with required reference video and optional text/image/audio references. [Routing] Use for cost-efficient video-to-video transformations. - [Seedream 5.0 Pro](https://docs.modellix.ai/bytedance/seedream-5-0-pro.md): [Core Function] Seedream 5.0 Pro is ByteDance's flagship professional-grade Text-to-Image (T2I) generation model. [Strengths] It delivers top-tier image quality with enhanced precision control over positions and elements, superior prompt adherence, and improved generation consistency for professiona… - [Seedream 5.0 Pro Edit](https://docs.modellix.ai/bytedance/seedream-5-0-pro-edit.md): [Core Function] Seedream 5.0 Pro Edit is a professional-grade single-image editing (I2I) model. [Strengths] It supports interactive precise editing: edit locations can be specified via coordinates, selection boxes, or arrows described in the prompt, with strong element-level control and subject cons… - [Seedream 5.0 Pro Multi Reference](https://docs.modellix.ai/bytedance/seedream-5-0-pro-multi-reference.md): [Core Function] Seedream 5.0 Pro Multi-Reference is a professional-grade multi-reference image generation (I2I) model that creates a single image from 2-10 reference images plus a text prompt. [Strengths] It excels at reference consistency, preserving characters, styles, and objects across multiple… - [Seedream 5.0 Lite](https://docs.modellix.ai/bytedance/seedream-5-0-lite.md): [Core Function] Seedream 5.0 Lite is a smart, reasoning-enhanced image generation model with real-time web search capabilities. [Strengths] It excels at generating time-sensitive imagery, infographics, and content requiring deep world knowledge or online search, boasting superior prompt understandin… - [Seedream 5.0 Lite Edit](https://docs.modellix.ai/bytedance/seedream-5-0-lite-edit.md): [Core Function] Seedream 5.0 Lite Edit is a reasoning-enhanced, smart image editing model. [Strengths] It features superior cross-modal understanding and reasoning, allowing for highly accurate, interactive multi-turn image editing with real-time knowledge enhancement. [Best For] Highly recommended… - [Nano Banana 2](https://docs.modellix.ai/google/nano-banana-2.md): [Core Function] Nano Banana 2 (Gemini 3.1 Flash Image) is an extremely fast text-to-image model. [Strengths] It is optimized for high-speed, high-volume visual generation, creative prompting, and rapid stylistic experimentation. [Best For] Highly recommended for: rapid creative iteration, generating… - [Nano Banana 2 Edit](https://docs.modellix.ai/google/nano-banana-2-edit.md): [Core Function] Nano Banana 2 Edit is a high-speed image editing model. [Strengths] It rapidly modifies existing images or extracts image frames from videos based on text prompts. [Best For] Highly recommended for: rapid style transfer, quick image modifications, and fast creative edits. [Limitation… - [Nano Banana 2 Lite](https://docs.modellix.ai/google/nano-banana-2-lite.md): [Core Function] Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is the lightweight, cost-efficient text-to-image model of the Nano Banana 2 family. [Strengths] It generates images even faster and more cheaply than Nano Banana 2, well suited to high-volume creative and stylized output at lower cost.… - [Nano Banana 2 Lite Edit](https://docs.modellix.ai/google/nano-banana-2-lite-edit.md): [Core Function] Nano Banana 2 Lite Edit (Gemini 3.1 Flash Lite Image) is the lightweight, cost-efficient image editing model of the Nano Banana 2 family; it transforms one or more input images per a text instruction. [Strengths] It performs fast, low-cost instruction-based editing across up to 14 in… - [Nano Banana Pro](https://docs.modellix.ai/google/nano-banana-pro.md): [Core Function] Nano Banana Pro (Gemini 3 Pro Image) is a high-capability creative image model. [Strengths] It balances the creative flexibility and speed of the Nano series with higher fidelity output. [Best For] Highly recommended for: high-quality stylized art and complex creative compositions. [… - [Nano Banana Pro Edit](https://docs.modellix.ai/google/nano-banana-pro-edit.md): [Core Function] Nano Banana Pro Edit is a high-capability creative image editing model. [Strengths] It provides high-quality creative edits, background replacements, and style transformations. [Best For] Highly recommended for: detailed creative modifications and complex style transfers. [Limitation… - [Nano Banana](https://docs.modellix.ai/google/nano-banana.md): [Core Function] Nano Banana is the original fast creative image model. [Strengths] Very fast creative generation. [Best For] Quick sketches and ideas. [Limitations] Superseded by Nano Banana 2 for general speed tasks. [Routing] Default to Nano Banana 2 unless specifically requested. - [Nano Banana Edit](https://docs.modellix.ai/google/nano-banana-edit.md): [Core Function] Nano Banana Edit is the original fast image editing model. [Strengths] Fast basic edits. [Limitations] Superseded by Nano Banana 2 Edit. [Routing] Default to Nano Banana 2 Edit unless specifically requested. - [Veo 3.1 T2V](https://docs.modellix.ai/google/veo-3-1-t2v.md): [Core Function] Veo 3.1 T2V is Google's state-of-the-art cinematic text-to-video engine. [Strengths] It natively generates 4K professional-grade video output with natively synchronized audio and supports complex camera movements. [Best For] Highly recommended for: high-end creative storytelling, cin… - [Veo 3.1 I2V](https://docs.modellix.ai/google/veo-3-1-i2v.md): [Core Function] Veo 3.1 I2V is Google's cinematic image-to-video generation model. [Strengths] It generates high-fidelity 4K video from a starting image. It supports advanced features like first-and-last frame conditioning and referencing up to three images. [Best For] Highly recommended for: animat… - [Veo 3.1 Fast T2V](https://docs.modellix.ai/google/veo-3-1-fast-t2v.md): [Core Function] Veo 3.1 Fast T2V is a high-speed text-to-video model. [Strengths] It is heavily optimized for fast generation, delivering video content with natively synchronized audio at 1080p quickly. [Best For] Highly recommended for: rapid prototyping, quick visual iteration, and high-volume bac… - [Veo 3.1 Fast I2V](https://docs.modellix.ai/google/veo-3-1-fast-i2v.md): [Core Function] Veo 3.1 Fast I2V is a high-speed image-to-video model. [Strengths] It quickly animates starting images at 1080p, optimized for low latency. [Best For] Highly recommended for: rapid prototyping and quick social media visual iterations. [Limitations] Do NOT use this model for the absol… - [Veo 3.1 Lite T2V](https://docs.modellix.ai/google/veo-3-1-lite-t2v.md): [Core Function] Veo 3.1 Lite T2V is a balanced text-to-video model. [Strengths] It provides a good balance between generation speed and visual quality, still supporting the advanced architecture of the 3.1 series. [Best For] Highly recommended for: general video content creation and social media pos… - [Veo 3.1 Lite I2V](https://docs.modellix.ai/google/veo-3-1-lite-i2v.md): [Core Function] Veo 3.1 Lite I2V is a balanced image-to-video model. [Strengths] It offers a middle ground between speed and quality for animating images. [Best For] Highly recommended for: general image animation and web-ready content. [Limitations] Do NOT use this model if you need 4K resolution.… - [Gemini Omni Flash T2V](https://docs.modellix.ai/google/gemini-omni-flash-t2v.md): [Core Function] Gemini Omni Flash T2V is Google's fast multimodal Text-to-Video generation model built on the Interactions API. [Strengths] It quickly turns a text prompt into a short 720p video with natively synchronized audio, offering low latency and solid prompt adherence. [Best For] Highly reco… - [Gemini Omni Flash I2V](https://docs.modellix.ai/google/gemini-omni-flash-i2v.md): [Core Function] Gemini Omni Flash I2V is a fast Image-to-Video model that animates a single input image into a short 720p video via the Interactions API. [Strengths] It uses the provided image as the opening frame and generates smooth motion with natively synchronized audio at low latency. [Best For… - [Gemini Omni Flash R2V](https://docs.modellix.ai/google/gemini-omni-flash-r2v.md): [Core Function] Gemini Omni Flash R2V (Reference-to-Video) generates a short 720p video guided by up to three reference images via the Interactions API. [Strengths] It fuses the styles, subjects, or elements from multiple reference images (referred to in the text prompt) into a single coherent anima… - [Gemini Omni Flash Video Edit](https://docs.modellix.ai/google/gemini-omni-flash-video-edit.md): [Core Function] Gemini Omni Flash Video Edit performs conversational, instruction-driven editing of an existing video via the Interactions API. [Strengths] It applies natural-language edits (changing the scene, mood, style, lighting, background, or time of day) to an input video while preserving the… - [Gemini Omni 1.1 Flash T2V](https://docs.modellix.ai/google/gemini-omni-1-1-flash-t2v.md): [Core Function] Gemini Omni 1.1 Flash T2V generates a short video with synchronized audio from a text prompt. [Strengths] It supports 360p, 720p, 1080p, and 4k output, 16:9 or 9:16, and 3 to 10 second clips with native speech, music, and sound effects. [Best For] Highly recommended for: rapid protot… - [Gemini Omni 1.1 Flash I2V](https://docs.modellix.ai/google/gemini-omni-1-1-flash-i2v.md): [Core Function] Gemini Omni 1.1 Flash I2V animates a single image into a short video with synchronized audio. [Strengths] It uses the input image as the opening frame and supports 360p, 720p, 1080p, and 4k output, 16:9 or 9:16, and 3 to 10 second clips. [Best For] Highly recommended for: animating a… - [Gemini Omni 1.1 Flash V2V](https://docs.modellix.ai/google/gemini-omni-1-1-flash-v2v.md): [Core Function] Gemini Omni 1.1 Flash V2V edits an existing video from a text instruction. [Strengths] It can change scene, mood, style, lighting, or time of day while keeping the source length and aspect ratio, with optional 360p to 4k output and synchronized audio. [Best For] Highly recommended fo… - [Gemini 3.1 Flash TTS](https://docs.modellix.ai/google/gemini-3-1-flash-tts.md): [Core Function] Gemini 3.1 Flash TTS is Google's low-latency, controllable text-to-speech model. [Strengths] Single-speaker and two-speaker dialogue, 30 prebuilt voices, 70+ languages via language_code, and expressive delivery through style prompts plus inline audio tags such as [whispers], [slow],… - [Gemini 3.5 Transcribe](https://docs.modellix.ai/google/gemini-3-5-transcribe.md): [Core Function] Gemini 3.5 Transcribe is Google's speech-to-text model for complete pre-recorded audio. [Strengths] Accurate multilingual recognition across 85+ languages with optional language hints, custom vocabulary for brand names, speaker labels (up to 8 speakers), word-level timestamps, and Sm… - [Kling Video O1 T2V](https://docs.modellix.ai/kling/kling-video-o1-t2v.md): [Core Function] Kling Video O1 T2V is the text-to-video slice of Kling O1 Omni Video. [Strengths] Reasoning-enhanced prompt planning with 3-10s duration and 720p/1080p output. [Best For] Complex physical interactions and logically demanding scenes from text alone. [Limitations] No native audio; dura… - [Kling Video O1 I2V](https://docs.modellix.ai/kling/kling-video-o1-i2v.md): [Core Function] Kling Video O1 I2V is the image-to-video slice of Kling O1 Omni Video. [Strengths] Reasoning-enhanced generation from 1-7 reference images, 720p/1080p, duration 3-10s (single image only 5 or 10). [Best For] Complex physical motion grounded in reference frames. [Limitations] No native… - [Kling Video O1](https://docs.modellix.ai/kling/kling-video-o1.md): [Core Function] Kling Video O1 is the world's first reasoning-enhanced video model. [Strengths] It performs deep planning over the prompt before generation, delivering best-in-class physical consistency, complex motion logic, and strict adherence to long-form semantics. [Best For] Highly recommended… - [Kling Image O1](https://docs.modellix.ai/kling/kling-image-o1.md): [Core Function] Kling Image O1 is a reasoning-enhanced multimodal image model. [Strengths] It performs deep reasoning over prompts and references to handle complex logic, spatial relationships, and intricate multi-image combinations. [Best For] Highly recommended for: complex scenes requiring strict… - [Kling V3 T2V](https://docs.modellix.ai/kling/kling-v3-t2v.md): [Core Function] Kling V3 T2V is the next-generation text-to-video base model. [Strengths] It natively supports generating ultra-long 15-second videos, 4K resolution, and synchronized native audio directly from text. [Best For] Highly recommended for: high-end cinematic creation, 4K video generation,… - [Kling V3 I2V](https://docs.modellix.ai/kling/kling-v3-i2v.md): [Core Function] Kling V3 I2V is the next-generation image-to-video model. [Strengths] It transforms static images into video with support for 4K resolution, 15-second durations, and native audio, providing superior motion and character expressiveness. [Best For] Highly recommended for: animating con… - [Kling V3 T2I](https://docs.modellix.ai/kling/kling-v3-t2i.md): [Core Function] Kling V3 T2I is the flagship text-to-image model (POST /images/generations, model_name=kling-v3). [Strengths] High aesthetic quality, prompt adherence, 1K/2K. [Best For] Concept art and photorealistic generation without a reference image. [Limitations] No reference image; for I2I use… - [Kling V3 I2I](https://docs.modellix.ai/kling/kling-v3-i2i.md): [Core Function] Kling V3 I2I is the flagship image-to-image editing model (POST /images/generations, model_name=kling-v3 with image). [Strengths] High-quality style transfer and editing up to 2K. [Best For] Single-reference image editing. [Limitations] Do NOT send negative_prompt when image is prese… - [Kling V3 Omni T2V](https://docs.modellix.ai/kling/kling-v3-omni-t2v.md): [Core Function] Kling V3 Omni T2V is a multimodal-leaning text-to-video model in the V3 family, oriented toward stronger semantic control and subject consistency in prompt-led generation. [Strengths] It targets high-fidelity cinematic clips with native audio options, flexible 3-15s duration, and bet… - [Kling V3 Omni I2V](https://docs.modellix.ai/kling/kling-v3-omni-i2v.md): [Core Function] Kling V3 Omni I2V is a multimodal image-to-video model that animates from one or more reference images with stronger subject and style consistency. [Strengths] It accepts an images array for reference-led motion, aiming to preserve identity, wardrobe, and product look across the clip… - [Kling V3 Omni Video](https://docs.modellix.ai/kling/kling-v3-omni-video.md): [Core Function] Kling V3 Omni Video V2V is a multimodal video-to-video endpoint that edits or restyles existing footage using prompt plus optional image and video references. [Strengths] It focuses on source fidelity and subject consistency for Omni-style edit workflows, combining prompt guidance wi… - [Kling V3 Omni Image](https://docs.modellix.ai/kling/kling-v3-omni-image.md): [Core Function] Kling V3 Omni Image is a unified multimodal image generation endpoint (POST /images/omni-image). [Strengths] Multi-image reference, up to 4K, and optional series generation via result_type/series_amount. Use <<>> placeholders in prompt. [Best For] Character consistency, fusi… - [Kling V3 Turbo T2V](https://docs.modellix.ai/kling/kling-v3-turbo-t2v.md): [Core Function] Kling V3 Turbo T2V is a speed- and cost-optimized text-to-video model in the V3 family for fast short-form generation. [Strengths] It emphasizes lower latency and efficient throughput with native audio and improved lip-sync for talking-head style clips, typically targeting practical… - [Kling V3 Turbo I2V](https://docs.modellix.ai/kling/kling-v3-turbo-i2v.md): [Core Function] Kling V3 Turbo I2V is a speed- and cost-optimized image-to-video model that animates a single keyframe into short motion clips. [Strengths] It prioritizes fast turnaround and efficient generation with optional native audio and strong lip-sync for portrait or product first-frame anima… - [Kling Video Effects](https://docs.modellix.ai/kling/kling-video-effects.md): [Core Function] Kling Video Effects applies predefined visual effects to images. [Strengths] It automatically transforms 1 or 2 images into engaging short videos using viral/predefined effect templates. [Best For] Highly recommended for: social media trends and quick visual gags. [Limitations] Do NO… - [Kling Avatar](https://docs.modellix.ai/kling/kling-avatar.md): [Core Function] Kling Avatar is a specialized portrait animation model. [Strengths] It precisely animates a portrait image to lip-sync with an audio file or TTS audio ID. [Best For] Highly recommended for: virtual presenters, talking head videos, and digital avatars. [Limitations] Do NOT use this mo… - [Kling Image Expansion](https://docs.modellix.ai/kling/kling-image-expansion.md): [Core Function] Kling Image Expansion is an outpainting model. [Strengths] It intelligently extends the borders of an image (horizontal, vertical, or asymmetric) while matching the original style and context. [Best For] Highly recommended for: changing aspect ratios, extending landscapes, and fillin… - [Kolors Virtual Try On V1-5](https://docs.modellix.ai/kling/kolors-virtual-try-on-v1-5.md): [Core Function] Kolors Virtual Try-On V1.5 is a specialized AI fashion model. [Strengths] It highly accurately applies garments (including Top+Bottom combinations) onto a person's image, preserving fabric texture and draping. [Best For] Highly recommended for: e-commerce virtual fitting rooms and fa… - [Kolors Virtual Try On V1](https://docs.modellix.ai/kling/kolors-virtual-try-on-v1.md): [Core Function] Kolors Virtual Try-On V1 is a legacy AI fashion and virtual try-on model. [Strengths] Maintained primarily for backward compatibility and stable API contracts. [Best For] Existing legacy e-commerce integrations that have not yet migrated to the newer try-on pipeline. [Limitations] Do… - [MAI Image 2.6](https://docs.modellix.ai/microsoft/mai-image-2-6.md): [Core Function] MAI Image 2.6 is Microsoft's latest text-to-image generation model in the MAI Image family. [Strengths] It improves text rendering, portraits, 3D imagery, and commercial photorealistic output compared with MAI Image 2.5. [Best For] Highly recommended for: marketing visuals, product h… - [MAI Image 2.6 Edit](https://docs.modellix.ai/microsoft/mai-image-2-6-edit.md): [Core Function] MAI Image 2.6 Edit is Microsoft's latest prompt-guided image editing model in the MAI Image family. [Strengths] It applies targeted edits to a single source image with the same quality gains as MAI Image 2.6 generation. [Best For] Highly recommended for: object edits, layout changes,… - [MAI Image 2.6 Flash](https://docs.modellix.ai/microsoft/mai-image-2-6-flash.md): [Core Function] MAI Image 2.6 Flash is the faster, lower-cost variant of MAI Image 2.6 text-to-image generation. [Strengths] It targets similar quality to MAI Image 2.6 with lower latency for high-throughput workloads. [Best For] Highly recommended for: production pipelines, batch generation, and la… - [MAI Image 2.6 Flash Edit](https://docs.modellix.ai/microsoft/mai-image-2-6-flash-edit.md): [Core Function] MAI Image 2.6 Flash Edit is the faster, lower-cost variant of MAI Image 2.6 image editing. [Strengths] It applies prompt-guided edits to a single source image with lower latency than MAI Image 2.6 Edit. [Best For] Highly recommended for: high-throughput edit APIs and production retou… - [MAI Transcribe 1.5](https://docs.modellix.ai/microsoft/mai-transcribe-1-5.md): [Core Function] MAI-Transcribe 1.5 transcribes a single public audio URL into text via an async task. [Strengths] Multi-lingual recognition, optional locale forcing, phrase-list biasing, and word-level timestamps. [Best For] Meeting notes, captions, and batch audio-to-text. [Limitations] Do NOT use… - [MiniMax H3 T2V](https://docs.modellix.ai/minimax/minimax-h3-t2v.md): [Core Function] MiniMax H3 T2V is a text-to-video generation model that creates video from a text prompt only. [Strengths] It supports 4-15 second clips, 768P or 2K output, and concrete aspect ratios from cinematic ultrawide to vertical. [Best For] Highly recommended for: prompt-only storyboards, ch… - [MiniMax H3 I2V](https://docs.modellix.ai/minimax/minimax-h3-i2v.md): [Core Function] MiniMax H3 I2V generates video from reference images plus a text prompt. [Strengths] It accepts up to 9 reference images and optional reference audios, with 4-15 second duration, 768P or 2K output, and optional aspect ratio defaulting to 16:9. [Best For] Highly recommended for: chara… - [MiniMax H3 FL2V](https://docs.modellix.ai/minimax/minimax-h3-fl2v.md): [Core Function] MiniMax H3 FL2V generates video guided by a first frame, a last frame, or both. [Strengths] It supports first-only, last-only, and first-plus-last conditioning with 4-15 second duration and 768P or 2K output. Output framing follows the input frame imagery. [Best For] Highly recommend… - [MiniMax H3 V2V](https://docs.modellix.ai/minimax/minimax-h3-v2v.md): [Core Function] MiniMax H3 V2V generates video guided by one or more reference videos plus a text prompt. [Strengths] It accepts up to 3 reference videos with optional reference images and audios, 4-15 second duration, 768P or 2K output, and optional aspect ratio defaulting to 16:9. [Best For] Highl… - [Hailuo 2.3 T2V](https://docs.modellix.ai/minimax/hailuo-2-3-t2v.md): [Core Function] Hailuo 2.3 T2V is a flagship text-to-video generation model optimized for human performance and stylization. [Strengths] It excels at capturing intricate human motion, nuanced facial micro-expressions, prompt adherence, and applying highly stylized aesthetics (e.g., anime, ink wash,… - [Hailuo 2.3 I2V](https://docs.modellix.ai/minimax/hailuo-2-3-i2v.md): [Core Function] Hailuo 2.3 I2V is a flagship image-to-video generation model optimized for character animation. [Strengths] It excels at animating human characters from a single image, maintaining consistent facial features, producing natural micro-expressions, and handling stylized artwork seamless… - [Hailuo 2.3 Fast I2V](https://docs.modellix.ai/minimax/hailuo-2-3-fast-i2v.md): [Core Function] Hailuo 2.3 Fast I2V is a high-speed, cost-effective image-to-video generation model. [Strengths] It excels at generating videos from images much faster and at roughly 50% lower cost than the standard 2.3 model, while still maintaining the 2.3 architecture's strength in human motion.… - [Hailuo 02 T2V](https://docs.modellix.ai/minimax/hailuo-02-t2v.md): [Core Function] Hailuo 02 T2V is a text-to-video generation model optimized for physical realism. [Strengths] It excels at complex physics simulation, fluid dynamics, broad cinematic scenes, and natively rendering 1080p video up to 10 seconds without downscaling. [Best For] Highly recommended for: p… - [Hailuo 02 I2V](https://docs.modellix.ai/minimax/hailuo-02-i2v.md): [Core Function] Hailuo 02 I2V is an image-to-video model optimized for physical realism and sustained high resolution. [Strengths] It excels at animating broad scenes, maintaining complex physics, and supporting native 1080p generation for up to 10 seconds. [Best For] Highly recommended for: animati… - [Hailuo 02 FL2V](https://docs.modellix.ai/minimax/hailuo-02-fl2v.md): [Core Function] Hailuo 02 FL2V is a First-Last frame transition video model. [Strengths] It excels at generating a logical, physically accurate video transition that bridges a provided starting frame and an ending frame. [Best For] Highly recommended for: visual morphing, before-and-after transition… - [Speech 2.8 HD](https://docs.modellix.ai/minimax/speech-2-8-hd.md): [Core Function] MiniMax Speech 2.8 HD is a high-quality text-to-speech model that converts text into natural spoken audio, including expressive paralinguistic cues such as (laughs) and (sighs). [Strengths] Strong narration quality, stable prosody controls (speed, volume, pitch, emotion), optional ti… - [Speech 2.8 Turbo](https://docs.modellix.ai/minimax/speech-2-8-turbo.md): [Core Function] MiniMax Speech 2.8 Turbo is a lower-latency text-to-speech model with the same control surface as Speech 2.8 HD, including paralinguistic tags such as (laughs). [Strengths] Faster and more cost-efficient synthesis while retaining prosody, timbre mix, pronunciation, subtitle, and audi… - [MiniMax Voice Clone](https://docs.modellix.ai/minimax/minimax-voice-clone.md): [Core Function] MiniMax Voice Clone clones a speaker from a public reference audio URL and synthesizes new speech in one async request; only the final audio is returned. [Strengths] No voice enrollment management; optional language_boost for clone and synthesis; prosody controls (speed, volume, pitc… - [GPT Image 2.5 Sunburst](https://docs.modellix.ai/openai/gpt-image-2-5-sunburst.md): [Core Function] GPT Image 2.5 Sunburst is a quality-oriented text-to-image generation model. [Strengths] Supports ten aspect ratios, resolution tiers up to 4K, and transparent backgrounds. [Best For] Landscape illustrations, product posters, and transparent icons. [Limitations] Do NOT use this for e… - [GPT Image 2.5 Sunburst Edit](https://docs.modellix.ai/openai/gpt-image-2-5-sunburst-edit.md): [Core Function] GPT Image 2.5 Sunburst Edit is a quality-oriented image editing model. [Strengths] Supports up to 16 input images, optional mask-based local edits, resolution tiers up to 4K, and transparent backgrounds. [Best For] Product image changes, masked object replacements, and edits guided b… - [GPT Image 2.5 Flare](https://docs.modellix.ai/openai/gpt-image-2-5-flare.md): [Core Function] GPT Image 2.5 Flare is a speed-oriented text-to-image generation model. [Strengths] Supports ten aspect ratios, resolution tiers up to 4K, and transparent backgrounds. [Best For] Landscape illustrations, product posters, and transparent icons. [Limitations] Do NOT use this for editin… - [GPT Image 2.5 Flare Edit](https://docs.modellix.ai/openai/gpt-image-2-5-flare-edit.md): [Core Function] GPT Image 2.5 Flare Edit is a speed-oriented image editing model. [Strengths] Supports up to 16 input images, optional mask-based local edits, resolution tiers up to 4K, and transparent backgrounds. [Best For] Product image changes, masked object replacements, and edits guided by mul… - [GPT Image 2](https://docs.modellix.ai/openai/gpt-image-2.md): [Core Function] GPT Image 2 is a state-of-the-art text-to-image generation model. [Strengths] It excels at generating highly detailed, photorealistic images from text descriptions, with native support for ultra-high resolutions including 2K and 4K (up to 3840x2160). [Best For] Highly recommended for… - [GPT Image 2 Edit](https://docs.modellix.ai/openai/gpt-image-2-edit.md): [Core Function] GPT Image 2 Edit is a high-resolution image-to-image editing model. [Strengths] It excels at making high-fidelity edits and style transformations to a single source image based on a text prompt, preserving details at up to 4K resolutions. [Best For] Highly recommended for: profession… - [Whisper 1](https://docs.modellix.ai/openai/whisper-1.md): [Core Function] OpenAI Whisper transcribes a single public audio URL into text via an async task. [Strengths] Multiple output formats (verbose_json with word/segment timestamps, plain text, SRT, VTT), optional language and prompt biasing. [Best For] Meeting notes, podcasts, captions, and batch audio… - [C1 T2V](https://docs.modellix.ai/pixverse/c1-t2v.md): [Core Function] PixVerse c1 T2V generates a video purely from a text prompt, with no input image. [Strengths] Strong prompt adherence and smooth motion; optional audio. [Best For] Turning an idea or script into video, concept visualization, story beats, social clips from text. [Limitations] Do NOT u… - [C1 I2V](https://docs.modellix.ai/pixverse/c1-i2v.md): [Core Function] PixVerse c1 I2V animates a single starting image into a video guided by a text prompt. [Strengths] Smooth, prompt-guided motion from one frame; optional audio. [Best For] Bringing a photo/illustration to life, product showcases, quick cinematic motion from a still. [Limitations] Do N… - [C1 FL2V](https://docs.modellix.ai/pixverse/c1-fl2v.md): [Core Function] PixVerse c1 first-last-frame generates a video that transitions from a start frame to an end frame, guided by a prompt. [Strengths] Controlled start/end composition with smooth interpolation. [Best For] Morphs, scene transitions, before/after motion. [Limitations] Do NOT use this if… - [C1 R2V](https://docs.modellix.ai/pixverse/c1-r2v.md): [Core Function] PixVerse c1 reference-to-video (fusion) generates a video from a prompt while preserving subjects from 1-7 reference images; each reference can be tagged as subject/background and named for @-reference in the prompt. [Strengths] Precise multi-subject composition and identity preserva… - [V6 T2V](https://docs.modellix.ai/pixverse/v6-t2v.md): [Core Function] PixVerse v6 T2V generates a video purely from a text prompt, with no input image. [Strengths] Strong prompt adherence and smooth motion; optional audio and multi-clip generation. [Best For] Turning an idea or script into video, concept visualization, story beats, social clips from te… - [V6 I2V](https://docs.modellix.ai/pixverse/v6-i2v.md): [Core Function] PixVerse v6 I2V animates a single starting image into a video guided by a text prompt. [Strengths] Smooth, prompt-guided motion from one frame; optional audio and multi-clip. [Best For] Bringing a photo/illustration to life, product showcases, quick cinematic motion from a still. [Li… - [V6 FL2V](https://docs.modellix.ai/pixverse/v6-fl2v.md): [Core Function] PixVerse v6 first-last-frame generates a video that transitions from a start frame to an end frame, guided by a prompt. [Strengths] Controlled start/end composition with smooth interpolation. [Best For] Morphs, scene transitions, before/after motion. [Limitations] Do NOT use this if… - [V6 R2V](https://docs.modellix.ai/pixverse/v6-r2v.md): [Core Function] PixVerse v6 reference-to-video (fusion) generates a video from a prompt while preserving subjects from 1-7 reference images; each reference can be tagged as subject/background and named for @-reference in the prompt. [Strengths] Precise multi-subject composition and identity preserva… - [V6 Video Extend](https://docs.modellix.ai/pixverse/v6-video-extend.md): [Core Function] PixVerse v6 Extend continues an existing video, generating additional seconds guided by a text prompt. [Strengths] Seamless continuation of the existing motion and scene. [Best For] Lengthening clips, continuing an action, adding an ending to footage. [Limitations] Do NOT use this to… - [Lipsync](https://docs.modellix.ai/pixverse/lipsync.md): [Core Function] PixVerse Lip Sync drives a talking video so the subject's lips match given audio or text-to-speech. [Strengths] Accurate lip synchronization for talking-head videos; supports either an existing audio track or TTS from a chosen speaker voice. [Best For] Dubbing, virtual presenters, ch… - [Motion Control](https://docs.modellix.ai/pixverse/motion-control.md): [Core Function] PixVerse Motion Control (Mimic) animates a subject image so it follows the motion of a reference video. [Strengths] Transfers human/animal motion from a driving video onto a still subject. [Best For] Making a character mimic a dance or action, motion retargeting onto a photo. [Limita… - [Upscale Video](https://docs.modellix.ai/pixverse/upscale-video.md): [Core Function] PixVerse Upscale increases the resolution and clarity of an existing video. [Strengths] Sharper detail and higher-resolution output without changing content. [Best For] Enhancing low-resolution footage, finalizing clips for delivery. [Limitations] Requires an input video; it enhances… - [Video Restyle](https://docs.modellix.ai/pixverse/video-restyle.md): [Core Function] PixVerse Restyle re-renders an existing video into a new visual style. [Strengths] Consistent style transfer across all frames. [Best For] Turning footage into anime/3D/painterly looks, stylized remixes. [Limitations] Do NOT use this to change content, motion, or add new scenes; it o… - [Vidu Q3 Pro T2V](https://docs.modellix.ai/vidu/viduq3-pro-t2v.md): [Core Function] Vidu Q3 Pro T2V is a premium cinematic text-to-video generation model. [Strengths] It excels at generating top-tier, realistic videos from text with support for advanced multi-shot 'smart cuts', complex physics, and simultaneous audio-visual generation. [Best For] Highly recommended… - [Vidu Q3 Pro I2V](https://docs.modellix.ai/vidu/viduq3-pro-i2v.md): [Core Function] Vidu Q3 Pro I2V is a premium Image-to-Video generation model. [Strengths] It excels at transforming a single starting image into high-fidelity, cinematic video with stable character consistency, complex motion, and synchronized audio-visual capabilities. [Best For] Highly recommended… - [Vidu Q3 Pro FL2V](https://docs.modellix.ai/vidu/viduq3-pro-fl2v.md): [Core Function] Vidu Q3 Pro FL2V is a premium First-Last frame transition video model. [Strengths] It excels at generating highly detailed, cinematic, and logically consistent video transitions between a starting image and an ending image. [Best For] Highly recommended for: high-end commercial trans… - [Vidu Q3 Pro Fast I2V](https://docs.modellix.ai/vidu/viduq3-pro-fast-i2v.md): [Core Function] Vidu Q3 Pro Fast I2V is a high-speed Image-to-Video generation model. [Strengths] It excels at generating smooth, physically accurate continuous motion from a single starting frame with extremely low latency. [Best For] Highly recommended for: fast prototyping, short dynamic product… - [Vidu Q3 Turbo T2V](https://docs.modellix.ai/vidu/viduq3-turbo-t2v.md): [Core Function] Vidu Q3 Turbo T2V is a fast text-to-video generation model. [Strengths] It excels at rapidly generating smooth, dynamic videos from text descriptions with very low latency. [Best For] Highly recommended for: fast prototyping, quick visual brainstorming, generating background b-roll,… - [Vidu Q3 Turbo I2V](https://docs.modellix.ai/vidu/viduq3-turbo-i2v.md): [Core Function] Vidu Q3 Turbo I2V is a fast Image-to-Video generation model. [Strengths] It excels at quickly animating a starting frame into a video sequence with strong motion dynamics and low latency. [Best For] Highly recommended for: rapid social media content creation, quick animatics, and fas… - [Vidu Q3 Turbo FL2V](https://docs.modellix.ai/vidu/viduq3-turbo-fl2v.md): [Core Function] Vidu Q3 Turbo FL2V is a fast First-Last frame transition video model. [Strengths] It excels at rapidly generating a smooth video transition bridging a specific starting image and an ending image. [Best For] Highly recommended for: quick visual morphs, before-and-after transitions, ti… - [Vidu Q3 Turbo R2V](https://docs.modellix.ai/vidu/viduq3-turbo-r2v.md): [Core Function] Vidu Q3 Turbo R2V is a fast reference-to-video generation model. [Strengths] It excels at quickly generating dynamic videos based on a text prompt while preserving the identity of the subjects from provided reference images. [Best For] Highly recommended for: rapid character animatio… - [Vidu Q3 R2V](https://docs.modellix.ai/vidu/viduq3-r2v.md): [Core Function] Vidu Q3 R2V is a high-quality reference-to-video generation model. [Strengths] It excels at generating detailed, cinematic videos that precisely follow a text prompt while highly preserving the character identity from provided reference images. [Best For] Highly recommended for: prof… - [Vidu Q3 Mix R2V](https://docs.modellix.ai/vidu/viduq3-mix-r2v.md): [Core Function] Vidu Q3 Mix R2V is a mixed-style reference-to-video generation model. [Strengths] It excels at generating highly consistent character videos by synthesizing and blending features from multiple reference images (up to 7) based on a text prompt. [Best For] Highly recommended for: maint… - [Vidu Q3 AD](https://docs.modellix.ai/vidu/viduq3-ad.md): [Core Function] Vidu Q3 AD is a keyframe-driven short-play (short drama) Image-to-Video model that turns a film-style script plus character, scene, and prop reference images into a complete multi-shot short video, automatically planning shots and compositing them in one pass. [Strengths] It excels a… - [Vidu Q3 Drama](https://docs.modellix.ai/vidu/viduq3-drama.md): [Core Function] Vidu Q3 Drama (Short Play) is a script-to-video model that turns a written script plus character, scene, and prop reference images into a complete multi-shot short drama, automatically planning the shots, transitions, and camera work in a single pass. [Strengths] It excels at multi-s… - [Vidu Q2 Pro Digital Human](https://docs.modellix.ai/vidu/viduq2-pro-digital-human.md): [Core Function] Vidu Q2 Pro Digital Human is a premium portrait animation model. [Strengths] It excels at generating highly realistic, expressive digital humans from a single portrait image, featuring precise lip-sync to audio and natural facial micro-expressions. [Best For] Highly recommended for:… - [Vidu Q2 Pro Multi Frame](https://docs.modellix.ai/vidu/viduq2-pro-multi-frame.md): [Core Function] Vidu Q2 Pro Multi-Frame is a sequence-based animation model. [Strengths] It excels at creating continuous, high-quality animation by interpolating through a provided sequence of keyframes (up to 9 images). [Best For] Highly recommended for: complex motion control, precise character p… - [Vidu Q2 Turbo Digital Human](https://docs.modellix.ai/vidu/viduq2-turbo-digital-human.md): [Core Function] Vidu Q2 Turbo Digital Human is a fast portrait animation model. [Strengths] It excels at quickly animating a static portrait image into a speaking or moving digital human, syncing lip movements to provided audio with low latency. [Best For] Highly recommended for: rapid generation of… - [Vidu Q2 Turbo Multi Frame](https://docs.modellix.ai/vidu/viduq2-turbo-multi-frame.md): Vidu Q2 Turbo multi-frame animation model. Animates through a sequence of up to 9 frames (1 start + up to 8 key images). Both `start_image` and `key_images` are required. Supports `resolution`. - [Template Story](https://docs.modellix.ai/vidu/template-story.md): [Core Function] Vidu Template Story is a narrative video generation model. [Strengths] It excels at placing user-provided character images into predefined, structured narrative templates (like 'love_story' or 'monkey_king') to automatically generate a cohesive short film. [Best For] Highly recommend… - [One Click AD Film](https://docs.modellix.ai/vidu/one-click-ad-film.md): [Core Function] Vidu One-Click AD-Film is an automated marketing video generation model. [Strengths] It excels at transforming 1 to 7 product or scene images into a polished, commercial-style advertisement video (10-60s) automatically. [Best For] Highly recommended for: e-commerce product showcases,… - [One Click General Film](https://docs.modellix.ai/vidu/one-click-general-film.md): [Core Function] Vidu One-Click General Film is an automated cinematic film generation model. [Strengths] It excels at automatically stringing together 1 to 7 user-provided images into a cohesive, cinematic film (up to 180s) with appropriate transitions and pacing. [Best For] Highly recommended for:… - [One Click Trending Replicate](https://docs.modellix.ai/vidu/one-click-trending-replicate.md): [Core Function] Vidu One-Click Trending Replicate is a viral video style cloning model. [Strengths] It excels at analyzing a trending or viral reference video and recreating its specific visual style, transitions, and pacing using the user's own provided subject images. [Best For] Highly recommended… - [Lip Sync](https://docs.modellix.ai/vidu/lip-sync.md): [Core Function] Vidu Lip Sync is a video-to-video audio synchronization model. [Strengths] It excels at reanimating lip movements in an existing video to precisely match a new replacement audio track, while preserving the original face identity. [Best For] Highly recommended for: dubbing videos into… - [Motion Sync](https://docs.modellix.ai/vidu/motion-sync.md): [Core Function] Vidu Motion Sync is a video-to-video motion transfer model. [Strengths] It excels at accurately extracting physical motion from a source video (e.g., a dancing person) and applying it to a target character image, preserving the target's identity. [Best For] Highly recommended for: cr… - [Grok Imagine Image](https://docs.modellix.ai/xai/grok-imagine-image.md): [Core Function] Grok Imagine Image is xAI's standard text-to-image generation model. [Strengths] It excels at quickly generating solid, visually appealing images from a text prompt across a wide range of aspect ratios. [Best For] Highly recommended for: rapid prototyping, social media content, and g… - [Grok Imagine Image Edit](https://docs.modellix.ai/xai/grok-imagine-image-edit.md): [Core Function] Grok Imagine Image Edit is xAI's standard image editing model. [Strengths] It excels at quickly applying prompt-guided edits and style changes to one or more source images. [Best For] Highly recommended for: fast restyling, quick variations, and lightweight image edits. [Limitations]… - [Grok Imagine Image 2.0](https://docs.modellix.ai/xai/grok-imagine-image-2-0.md): [Core Function] Grok Imagine Image 2.0 generates images from a text prompt. [Strengths] It supports quality (low or medium), 1k or 2k resolution, a wide range of aspect ratios including 21:9 and 5:2, and up to 10 images per request. [Best For] Highly recommended for: concept art, marketing visuals,… - [Grok Imagine Image 2.0 Edit](https://docs.modellix.ai/xai/grok-imagine-image-2-0-edit.md): [Core Function] Grok Imagine Image 2.0 Edit edits one to three source images from a text prompt. [Strengths] It supports quality (low or medium), 1k or 2k resolution, and the same aspect ratios as Grok Imagine Image 2.0. [Best For] Highly recommended for: restyling, combining up to 3 references, and… - [Grok Imagine Image Quality](https://docs.modellix.ai/xai/grok-imagine-image-quality.md): [Core Function] Grok Imagine Image (Quality) is xAI's high-fidelity text-to-image generation model. [Strengths] It excels at producing richly detailed, high-quality images from a text prompt, with flexible aspect ratios and an optional 2K resolution. [Best For] Highly recommended for: detailed conce… - [Grok Imagine Image Quality Edit](https://docs.modellix.ai/xai/grok-imagine-image-quality-edit.md): [Core Function] Grok Imagine Image Edit (Quality) is xAI's high-fidelity image editing model. [Strengths] It excels at applying detailed, prompt-guided edits and style transformations to one or more source images while preserving fine detail. [Best For] Highly recommended for: high-quality restyling… - [Grok Imagine Video 1.5 I2V](https://docs.modellix.ai/xai/grok-imagine-video-1-5-i2v.md): [Core Function] Grok Imagine Video 1.5 I2V animates a single starting image into a video using the Grok Imagine 1.5 generation backbone. [Strengths] It excels at producing motion from one starting frame with the improved 1.5 model. [Best For] Highly recommended for: animating a photo or illustration… - [Grok Imagine Video](https://docs.modellix.ai/xai/grok-imagine-video.md): [Core Function] Grok Imagine Video is xAI's text-to-video generation model. [Strengths] It excels at generating short, dynamic video clips directly from a text prompt, with controllable duration, aspect ratio, and resolution. [Best For] Highly recommended for: short social clips, animated concepts,… - [Grok Imagine Video I2V](https://docs.modellix.ai/xai/grok-imagine-video-i2v.md): [Core Function] Grok Imagine Video I2V animates a single starting image into a video. [Strengths] It excels at producing smooth motion from one starting frame, guided by a text prompt for the desired movement. [Best For] Highly recommended for: bringing a photo or illustration to life, dynamic produ… - [Grok Imagine Video R2V](https://docs.modellix.ai/xai/grok-imagine-video-r2v.md): [Core Function] Grok Imagine Video R2V generates a video from a text prompt while preserving the subjects shown in up to 7 reference images. [Strengths] It excels at keeping character/subject identity consistent across a newly generated scene driven by the prompt. [Best For] Highly recommended for:… - [Grok Imagine Video Edit](https://docs.modellix.ai/xai/grok-imagine-video-edit.md): [Core Function] Grok Imagine Video Edit applies a prompt-guided transformation to an input video. [Strengths] It excels at restyling and modifying an existing video while keeping its original duration and aspect ratio. [Best For] Highly recommended for: restyling clips, applying visual effects, and… - [Grok Imagine Video Extend](https://docs.modellix.ai/xai/grok-imagine-video-extend.md): [Core Function] Grok Imagine Video Extend continues an existing video, generating additional footage beyond its end. [Strengths] It excels at seamlessly extending a clip with new prompt-guided motion. [Best For] Highly recommended for: lengthening short clips, continuing a scene, and adding follow-o… - [Grok Voice TTS](https://docs.modellix.ai/xai/grok-voice-tts.md): [Core Function] Grok Voice TTS converts text into natural spoken audio with expressive voices and optional speech tags embedded in the text. [Strengths] Supports 20+ languages (plus auto-detect), 26 built-in voices, multiple codecs (mp3/wav/pcm/mulaw/alaw), and custom voice IDs. [Best For] Product v… - [Grok Voice ASR](https://docs.modellix.ai/xai/grok-voice-asr.md): [Core Function] Grok Voice ASR transcribes a single public audio URL into text via an async task. [Strengths] Word-level timestamps, optional speaker diarization, multichannel transcription, Inverse Text Normalization (format + language), keyterm biasing, and filler-word control. [Best For] Meeting… - [Query Task Result](https://docs.modellix.ai/api/get-task-result.md): Query the status and results of an async task by task_id - [Upload Media File](https://docs.modellix.ai/api/upload-media-file.md): Upload a single media file via `multipart/form-data` (field name: `file`). Returns a `file_id` and `url` you can pass into prediction APIs. See [Upload media files](/ways-to-use/api#upload-media-files) for limits, supported formats, and the full workflow. - [List Media Files](https://docs.modellix.ai/api/list-media-files.md): List non-expired media files belonging to the authenticated team. Default page size is 100; `limit` is capped at 100. See [Upload media files](/ways-to-use/api#upload-media-files) for the full workflow. - [Delete Media File](https://docs.modellix.ai/api/delete-media-file.md): Delete a media file. After deletion, the file no longer counts toward your upload limit. Returns `404` if the file is not found. See [Upload media files](/ways-to-use/api#upload-media-files) for the full workflow. - [List Active Models](https://docs.modellix.ai/api/list-models.md): Returns currently active (published) models with slug, type, name, series name, documentation URL, description, and optional display price. Optional query `featured=true` limits results to CMS featured models. - [Get Schema](https://docs.modellix.ai/api/get-schema.md): Returns the request and response schema for a Media Model. Pass the model slug in the path (for example alibaba/qwen-image-3.0-pro). This endpoint is public and does not require an API key. The response includes servers (inference base URL) and post (OpenAPI-style operation: description, requestBody… - [Get Logs](https://docs.modellix.ai/api/get-logs.md): Returns paginated media model request logs for the team that owns the API Key. Requires a time window of at most 30 days. Optional `mdlx_user_id` filters by the end-user id sent as `X-Mdlx-User-Id` on async inference. Subject to the query rate limit (separate from media generation RPM); responses ma… - [STT Result Schema](https://docs.modellix.ai/media-model-api/stt-result-schema.md): Normalized STT result JSON (modellix.transcript.v1) from document resources—full text, channels, sentence and word timings, and optional speaker diarization. - [Modellix LLM Overview](https://docs.modellix.ai/llm/overview.md): Get started with the Modellix LLM gateway—OpenAI-compatible Chat Completions and Responses, Anthropic-compatible Messages, and guides for SDKs and coding tools. - [Modellix LLM API Guide](https://docs.modellix.ai/llm/api/api.md): Call Modellix LLM with OpenAI-compatible Chat Completions and Responses or Anthropic-compatible Messages—sync requests, streaming SSE, multimodal content parts, auth, model list, request logs, errors, and billing. - [Chat Completions](https://docs.modellix.ai/llm/chat-completions.md): OpenAI Chat Completions–compatible endpoint on the Modellix LLM gateway. Accepts a synchronous request with `model` (provider/name format) and an OpenAI-style `messages` array; returns a chat.completion JSON object by default, or SSE (`text/event-stream`) when `stream=true`. Use this path for the Op… - [Responses](https://docs.modellix.ai/llm/responses.md): OpenAI Responses–compatible endpoint on the Modellix LLM gateway. Accepts a synchronous request with `model` (provider/name format) and `input` (string or content array); returns a Responses JSON object by default, or SSE (`text/event-stream`) when `stream=true`. Use `max_output_tokens` for output l… - [Messages](https://docs.modellix.ai/llm/messages.md): Anthropic Messages–compatible endpoint on the Modellix LLM gateway. Accepts a synchronous request with `model` (`provider/name`, use `anthropic/...`), Anthropic-style `messages`, and required `max_tokens`; optional `system` prompt is supported. Returns a Messages JSON object by default, or SSE (`tex… - [List models](https://docs.modellix.ai/llm/list-models.md): Returns an OpenAI-compatible list of models available on the Modellix LLM gateway. Each `data[].id` uses `provider/name` form. Each item also includes `name`, `provider_name`, and `series_name`. Uses the query rate limit (separate from inference RPM). - [Get LLM logs](https://docs.modellix.ai/llm/get-llm-logs.md): Lists LLM request logs for the authenticated team within a time window. Optional `mdlx_user_id` filters by the end-user id sent as `X-Mdlx-User-Id` on inference. Uses the query rate limit (separate from inference RPM). - [Use Modellix LLM with Claude Code](https://docs.modellix.ai/llm/agent/claude-code.md): Configure Claude Code to call Modellix with ANTHROPIC_BASE_URL (no /v1), a Modellix API key, and anthropic/... models. - [Use Modellix LLM with CodeBuddy](https://docs.modellix.ai/llm/agent/codebuddy.md): Add Modellix as a CodeBuddy custom model with an OpenAI-compatible endpoint, models.json, and provider/name model IDs. - [Use Modellix LLM with Codex CLI](https://docs.modellix.ai/llm/agent/codex.md): Configure Codex to use Modellix by setting OPENAI_API_KEY and openai_base_url in ~/.codex/config.toml with openai/... models. - [Use Modellix LLM with DeepSeek Harness](https://docs.modellix.ai/llm/agent/deepseek-harness.md): Connect DeepSeek Harness to Modellix LLM—install the dsh-modellix plugin, or add a custom llm-pi-ai provider in the Web UI or $DSH_HOME/settings.yaml. - [Use Modellix LLM with Grok Build](https://docs.modellix.ai/llm/agent/grok-build.md): Point Grok Build at the Modellix LLM gateway with custom models in ~/.grok/config.toml, MODELLIX_API_KEY, and provider/name model IDs. - [Use Modellix LLM with Hermes Agent](https://docs.modellix.ai/llm/agent/hermes.md): Point Hermes Agent at the Modellix LLM gateway with a Custom Endpoint, config.yaml provider custom, and provider/name model IDs. - [Use Modellix LLM with Junie CLI](https://docs.modellix.ai/llm/agent/junie.md): Add Modellix as a Junie custom LLM profile with OpenAICompletion, a full chat completions URL, and provider/name model IDs. - [Use Modellix LLM with Kilo Code](https://docs.modellix.ai/llm/agent/kilo-code.md): Add Modellix as a Kilo Code custom OpenAI Compatible provider with kilo.jsonc, baseURL, and provider/name model IDs. - [Use Modellix LLM with OpenClaw 🦞](https://docs.modellix.ai/llm/agent/openclaw.md): Add Modellix as a custom OpenAI-compatible provider in OpenClaw with models.providers, openai-completions, and provider/name model IDs. - [Use Modellix LLM with OpenCode](https://docs.modellix.ai/llm/agent/opencode.md): Add Modellix as a custom OpenAI-compatible provider in OpenCode with @ai-sdk/openai-compatible, baseURL, and provider/name model IDs. - [Use Modellix LLM with Pi Agent](https://docs.modellix.ai/llm/agent/pi.md): Add Modellix as a custom OpenAI-compatible provider in Pi with models.json, openai-completions, and provider/name model IDs. - [Use Modellix LLM with Qwen Code](https://docs.modellix.ai/llm/agent/qwen-code.md): Point Qwen Code at the Modellix LLM gateway with OPENAI_BASE_URL, OPENAI_API_KEY, and provider/name model IDs. - [Use Modellix LLM with Reasonix](https://docs.modellix.ai/llm/agent/reasonix.md): Point Reasonix at the Modellix LLM gateway with a custom OpenAI-compatible [[providers]] entry, MODELLIX_API_KEY in ~/.reasonix/.env, and provider/name model IDs. - [Use Modellix LLM with WorkBuddy](https://docs.modellix.ai/llm/agent/workbuddy.md): Add Modellix as a WorkBuddy custom OpenAI-compatible model with Endpoint, models.json, and provider/name model IDs. - [Use Modellix LLM with the OpenAI SDK](https://docs.modellix.ai/llm/sdk/openai-sdk.md): Point the OpenAI SDK at the Modellix LLM gateway with OPENAI_BASE_URL and your Modellix API key, then call Chat Completions or Responses with provider/name models. - [Use Modellix LLM with the Anthropic SDK](https://docs.modellix.ai/llm/sdk/anthropic-sdk.md): Point the Anthropic SDK at Modellix with ANTHROPIC_BASE_URL (no /v1) and your Modellix API key, then call Messages with anthropic/... models. - [Use Modellix LLM with the OpenAI Agents SDK](https://docs.modellix.ai/llm/sdk/openai-agents.md): Point the OpenAI Agents SDK at Modellix with AsyncOpenAI base_url, Chat Completions, and provider/name model IDs. - [Use Modellix LLM with the Claude Agent SDK](https://docs.modellix.ai/llm/sdk/claude-agent-sdk.md): Point the Claude Agent SDK at Modellix with ANTHROPIC_BASE_URL (no /v1), a Modellix API key, and anthropic/... model IDs. - [Use Modellix LLM with the Vercel AI SDK](https://docs.modellix.ai/llm/sdk/vercel-ai-sdk.md): Point the Vercel AI SDK at Modellix with createOpenAICompatible, baseURL, and provider/name model IDs. - [Use Modellix LLM with LangChain](https://docs.modellix.ai/llm/framework/langchain.md): Point LangChain ChatOpenAI at the Modellix LLM gateway with base_url, a Modellix API key, and provider/name model IDs. - [Use Modellix LLM with Mastra](https://docs.modellix.ai/llm/framework/mastra.md): Point Mastra agents at Modellix with a custom OpenAI-compatible model url, apiKey, and gateway-style provider/name IDs. - [Use Modellix LLM with CrewAI](https://docs.modellix.ai/llm/framework/crewai.md): Point CrewAI at Modellix with LLM base_url, a Modellix API key, provider/name model IDs, and custom_openai for OpenAI-compatible routing. - [Use Modellix LLM with Agno](https://docs.modellix.ai/llm/framework/agno.md): Point Agno agents at Modellix with OpenAILike base_url, a Modellix API key, and provider/name model IDs. - [Use Modellix LLM with Microsoft Agent Framework](https://docs.modellix.ai/llm/framework/microsoft-agent-framework.md): Point Microsoft Agent Framework at Modellix with OpenAIChatCompletionClient base_url, a Modellix API key, and provider/name model IDs. - [Use Modellix LLM with Cursor](https://docs.modellix.ai/llm/ide/cursor.md): Add Modellix as an OpenAI-compatible provider in Cursor using https://llm.modellix.ai/v1, your Modellix API key, and provider/name model IDs. - [Use Modellix LLM with Cline](https://docs.modellix.ai/llm/ide/cline.md): Point Cline at the Modellix LLM gateway with the OpenAI Compatible provider, Base URL, and provider/name model IDs. - [Use Modellix LLM with CC Switch](https://docs.modellix.ai/llm/tool/cc-switch.md): Add Modellix as a custom CC Switch provider for Claude Code, Codex, and OpenAI Compatible apps with the correct Base URL per protocol. - [Modellix Tools Overview](https://docs.modellix.ai/tools/overview.md): Get started with Modellix Tools—Web Search and Web Fetch on https://tool.modellix.ai, including per-request and per-URL pricing. - [Install the Web Tools MCP](https://docs.modellix.ai/tools/mcp.md): Connect Cursor, VS Code, Claude Code, Codex, and other MCP clients to Modellix Web Search and Web Fetch at https://tool.modellix.ai/mcp. - [Web Search](https://docs.modellix.ai/api/web-search.md): Search the public web and return ranked results. - [Web Fetch](https://docs.modellix.ai/api/web-fetch.md): Extract readable content from public URLs. - [Get Tool Logs](https://docs.modellix.ai/api/get-tool-logs.md): Lists Web Search and Web Fetch request logs for the authenticated team within a time window. Optional `tool` filters by web-search or web-fetch. Uses the query rate limit (separate from tool call RPM). - [Modellix Product Updates and Announcements](https://docs.modellix.ai/changelog/product-updates.md): Follow Modellix product updates and announcements—new API endpoints, console features, billing changes, and platform improvements released each month. - [New AI Models Added to Modellix](https://docs.modellix.ai/changelog/new-models.md): Track newly integrated AI image, video, speech, and LLM models on Modellix, including provider, capabilities, parameters, and availability as they ship. - [Deprecated Models on Modellix](https://docs.modellix.ai/changelog/deprecated-models.md): See which AI image, video, and other media models are no longer available on Modellix, including the provider, model ID, and the date they were retired. ## OpenAPI Specs - [list-active-models](/else-api/list-active-models.json) - [media-files](/file-api/media-files.json) - [llm](/llm/api/llm.json) - [alibaba-i2i](/media-model-api/alibaba/alibaba-i2i.json) - [alibaba-i2v](/media-model-api/alibaba/alibaba-i2v.json) - [alibaba-s2s](/media-model-api/alibaba/alibaba-s2s.json) - [alibaba-s2t](/media-model-api/alibaba/alibaba-s2t.json) - [alibaba-t2i](/media-model-api/alibaba/alibaba-t2i.json) - [alibaba-t2s](/media-model-api/alibaba/alibaba-t2s.json) - [alibaba-t2v](/media-model-api/alibaba/alibaba-t2v.json) - [alibaba-v2v](/media-model-api/alibaba/alibaba-v2v.json) - [bytedance-i2i](/media-model-api/bytedance/bytedance-i2i.json) - [bytedance-i2v](/media-model-api/bytedance/bytedance-i2v.json) - [bytedance-t2i](/media-model-api/bytedance/bytedance-t2i.json) - [bytedance-t2v](/media-model-api/bytedance/bytedance-t2v.json) - [bytedance-v2v](/media-model-api/bytedance/bytedance-v2v.json) - [get-schema](/media-model-api/get-schema.json) - [google-i2i](/media-model-api/google/google-i2i.json) - [google-i2v](/media-model-api/google/google-i2v.json) - [google-s2t](/media-model-api/google/google-s2t.json) - [google-t2i](/media-model-api/google/google-t2i.json) - [google-t2s](/media-model-api/google/google-t2s.json) - [google-t2v](/media-model-api/google/google-t2v.json) - [google-v2v](/media-model-api/google/google-v2v.json) - [kling-i2i](/media-model-api/kling/kling-i2i.json) - [kling-i2v](/media-model-api/kling/kling-i2v.json) - [kling-t2i](/media-model-api/kling/kling-t2i.json) - [kling-t2v](/media-model-api/kling/kling-t2v.json) - [kling-v2v](/media-model-api/kling/kling-v2v.json) - [logs](/media-model-api/logs.json) - [microsoft-i2i](/media-model-api/microsoft/microsoft-i2i.json) - [microsoft-s2t](/media-model-api/microsoft/microsoft-s2t.json) - [microsoft-t2i](/media-model-api/microsoft/microsoft-t2i.json) - [minimax-i2i](/media-model-api/minimax/minimax-i2i.json) - [minimax-i2v](/media-model-api/minimax/minimax-i2v.json) - [minimax-s2s](/media-model-api/minimax/minimax-s2s.json) - [minimax-t2i](/media-model-api/minimax/minimax-t2i.json) - [minimax-t2s](/media-model-api/minimax/minimax-t2s.json) - [minimax-t2v](/media-model-api/minimax/minimax-t2v.json) - [minimax-v2v](/media-model-api/minimax/minimax-v2v.json) - [openai-i2i](/media-model-api/openai/openai-i2i.json) - [openai-s2t](/media-model-api/openai/openai-s2t.json) - [openai-t2i](/media-model-api/openai/openai-t2i.json) - [pixverse-i2v](/media-model-api/pixverse/pixverse-i2v.json) - [pixverse-t2v](/media-model-api/pixverse/pixverse-t2v.json) - [pixverse-v2v](/media-model-api/pixverse/pixverse-v2v.json) - [query-task-result](/media-model-api/query-task-result.json) - [reve-i2i](/media-model-api/reve/reve-i2i.json) - [reve-t2i](/media-model-api/reve/reve-t2i.json) - [skyreels-i2v](/media-model-api/skywork/skyreels-i2v.json) - [skyreels-t2v](/media-model-api/skywork/skyreels-t2v.json) - [skyreels-v2v](/media-model-api/skywork/skyreels-v2v.json) - [vidu-i2v](/media-model-api/vidu/vidu-i2v.json) - [vidu-t2v](/media-model-api/vidu/vidu-t2v.json) - [vidu-v2v](/media-model-api/vidu/vidu-v2v.json) - [xai-i2i](/media-model-api/xai/xai-i2i.json) - [xai-i2v](/media-model-api/xai/xai-i2v.json) - [xai-s2t](/media-model-api/xai/xai-s2t.json) - [xai-t2i](/media-model-api/xai/xai-t2i.json) - [xai-t2v](/media-model-api/xai/xai-t2v.json) - [xai-tts](/media-model-api/xai/xai-tts.json) - [xai-v2v](/media-model-api/xai/xai-v2v.json) - [get-team-balance](/team-api/get-team-balance.json) - [validate-api-key](/team-api/validate-api-key.json) - [logs](/tools/logs.json) - [web-tools](/tools/web-tools.json) ## Optional - [Support](mailto:support@modellix.ai) - [Community](https://discord.gg/N2FbcB2cZT) - [Blog](https://www.modellix.ai/blog/)