[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"docs-nav-en":3,"docs-nav-zh":933,"docs-page-zh-image\u002Fbytedance\u002Fseedream-5-0":1462},{"locale":4,"updatedAt":5,"items":6},"en","2026-09-20T05:24:50.614336931Z",[7,44,332,471,830,884,920],{"id":8,"groupId":9,"locale":4,"slug":10,"title":11,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"topTab":17,"children":18},"01a06b0c-7659-7fc4-9145-1bff398609e5","grp-section-quick-start","","Quick Start","overview","section",null,0,"LISTED","api",[19,28,36],{"id":20,"groupId":21,"locale":4,"slug":22,"title":23,"pageType":12,"contentSource":24,"contentRef":25,"sort":15,"status":16,"icon":12,"description":26,"hasToc":27},"01a06b0c-7cf6-75e2-b8e9-b089f7ec3c6d","grp-integration-overview","integration\u002Foverview","Overview","admin","01a06b0d-a30f-748e-99ad-34cebcf5a108","iCreat AI is an all-in-one generative AI model platform",true,{"id":29,"groupId":30,"locale":4,"slug":31,"title":11,"pageType":32,"contentSource":24,"contentRef":33,"sort":15,"status":16,"icon":34,"description":35,"hasToc":27},"01a06b0c-7dd5-7fc1-857d-cc97ab9bc5cf","grp-quick-start","quick-start","article","01a06b0d-adb3-7c50-9ae3-7d9214150386","quick","Get an API key and complete your first API call in 2 minutes",{"id":37,"groupId":38,"locale":4,"slug":39,"title":40,"pageType":32,"contentSource":24,"contentRef":41,"sort":15,"status":16,"icon":42,"description":43,"hasToc":27},"01a06b0c-7eac-75ec-8aa9-06d21bda4c47","grp-api-key","api-key","API Key","01a06b0d-b78f-7aca-a8d8-4b4da99ab051","key","Create and manage API keys for authentication",{"id":45,"groupId":46,"locale":4,"slug":10,"title":47,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"topTab":17,"children":48},"01a06b0c-7772-7fad-af6a-accca5f3e6e1","grp-section-llm","LLM",[49,114,151,178,206,227,262,283,297,318],{"id":50,"groupId":51,"locale":4,"slug":10,"title":52,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":53,"description":54,"hasToc":55,"children":56},"01a06b69-ba7c-7864-9e50-3502712dbf9b","grp-section-claude","Claude","claude","Designed for real-world enterprise agentic workflows, the Claude 5 lineup introduces a clear tiered intelligence architecture: Claude Fable 5 (the state-of-the-art flagship for complex long-horizon research), Claude Opus 5 (built for agentic software development and complex engineering tasks), Claude Sonnet 5 (the ultimate combination of speed and high intelligence), and Claude Haiku 4.5 (ultra-fast, cost-effective execution).",false,[57,65,72,79,86,93,100,107],{"id":58,"groupId":59,"locale":4,"slug":60,"title":61,"pageType":17,"contentSource":24,"contentRef":62,"sort":15,"status":16,"icon":53,"description":63,"httpMethod":64,"hasToc":27},"01a08015-a879-7697-84f7-d86878770080","grp-api-claude-fable-5-1","llm\u002Fclaude-fable-5-1","Claude Fable 5.1 Economy","01a08017-4a18-7241-9a31-2b86c7cdf351","Claude Fable 5.1 is Anthropic's upgraded flagship frontier model engineered for high-complexity software development, long-horizon agentic workflows, and multi-step knowledge work. Built on Mythos-tier reasoning, Fable 5.1 significantly enhances autonomous self-verification and root-cause troubleshooting while reducing safety false-positive interventions by up to 60%. Powered by an optimized prompt-caching architecture featuring a 75% reduction in cache-read pricing, it cuts total operating costs by 25% to 45% for highly agentic tasks.","POST",{"id":66,"groupId":67,"locale":4,"slug":68,"title":69,"pageType":17,"contentSource":24,"contentRef":70,"sort":15,"status":16,"icon":53,"description":71,"httpMethod":64,"hasToc":27},"01a06b69-e1f3-7007-9249-9fb507a091d6","grp-api-claude-fable-5","llm\u002Fclaude-fable-5","Claude Fable 5 Economy","01a06b6a-8a43-71f2-9a8f-8a781c7626f1","Claude Fable 5 is Anthropic’s officially released Mythos-tier model, designed for autonomous knowledge work, advanced reasoning, and long-horizon coding. It supports text, image, and file inputs and produces text output. It features a 1-million-token context window, up to 128,000 output tokens, and built-in always-on adaptive thinking, tool use, and structured outputs. The model shares the same underlying model family as Claude Mythos 5 but was released publicly with stronger safety safeguards. It is suitable for complex software engineering, multi-step agentic workflows, enterprise research, document analysis, and professional tasks requiring sustained reasoning and high reliability.",{"id":73,"groupId":74,"locale":4,"slug":75,"title":76,"pageType":17,"contentSource":24,"contentRef":77,"sort":15,"status":16,"icon":53,"description":78,"httpMethod":64,"hasToc":27},"01a06b6a-03d4-73da-bb15-dd4f1d337092","grp-api-claude-sonnet-5","llm\u002Fclaude-sonnet-5","Claude Sonnet 5 Economy","01a06b6b-56c7-720a-9c41-7289886812cb","Sonnet 5 is Anthropic’s most capable Sonnet-class model, delivering frontier-level performance across coding, agentic workflows, and professional tasks. It features adaptive thinking with selectable reasoning levels—**low**, **medium**, **high**, **max**, and **x-high**—along with a 1-million-token context window and support for text, image, and file inputs.\n\nBuilt with an updated tokenizer, Sonnet 5 also incorporates real-time cybersecurity safeguards designed to block high-risk dual-use activities.",{"id":80,"groupId":81,"locale":4,"slug":82,"title":83,"pageType":17,"contentSource":24,"contentRef":84,"sort":15,"status":16,"icon":53,"description":85,"httpMethod":64,"hasToc":27},"01a08015-a7a0-7c92-842d-29c1e25fd5e4","grp-api-claude-opus-5","llm\u002Fclaude-opus-5","Claude Opus 5 Economy","01a08017-44df-79e8-b97f-6bd35f1e3f5c","Claude Opus 5 is Anthropic's ultra-flagship frontier model engineered for extreme cognitive density, deep logical reasoning, and expert-level scientific synthesis. Powered by native adaptive thinking and a 1-million-token context window, it resolves complex multi-step mathematical, architectural, and coding challenges with unprecedented accuracy and near-zero hallucination rates. Coupled with rigorous enterprise safety alignment, Claude Opus 5 excels at high-stakes operations, including legal compliance, quantitative financial modeling, advanced biosecurity research, and mission-critical software architecture auditing.",{"id":87,"groupId":88,"locale":4,"slug":89,"title":90,"pageType":17,"contentSource":24,"contentRef":91,"sort":15,"status":16,"icon":53,"description":92,"httpMethod":64,"hasToc":27},"01a06b6a-04a8-7ab9-b234-cd683eff54c8","grp-api-claude-opus-4-8","llm\u002Fclaude-opus-4-8","Claude Opus 4.8 Economy","01a06b6b-5c21-792f-ba15-f13161d13758","Claude Opus 4.8 is Anthropic’s most powerful officially released Opus model, designed for complex reasoning, long-horizon agentic coding, and highly autonomous professional workflows. It supports text, image, and file inputs and produces text output. It features a 1-million-token context window, up to 128,000 output tokens, and built-in capabilities for adaptive thinking, tool use, and structured outputs. The model excels at advanced coding, browser and computer-use agents, enterprise knowledge work, financial and legal analysis, and multi-step tasks requiring sustained judgment and high reliability.",{"id":94,"groupId":95,"locale":4,"slug":96,"title":97,"pageType":17,"contentSource":24,"contentRef":98,"sort":15,"status":16,"icon":53,"description":99,"httpMethod":64,"hasToc":27},"01a06b6a-064d-746c-a7db-83fba6e3afe5","grp-api-claude-opus-4-7","llm\u002Fclaude-opus-4-7","Claude Opus 4.7 Economy","01a06b6b-668a-7850-a2fc-8df87f9551ea","Claude Opus 4.7 is Anthropic’s most powerful officially released Opus model, designed for complex reasoning, long-horizon agentic coding, and highly autonomous professional workflows. It supports text, image, and file inputs and produces text output. It features a 1-million-token context window, up to 128,000 output tokens, and built-in capabilities for adaptive thinking, tool use, and structured outputs. The model excels at advanced coding, browser and computer-use agents, enterprise knowledge work, financial and legal analysis, and multi-step tasks requiring sustained judgment and high reliability.",{"id":101,"groupId":102,"locale":4,"slug":103,"title":104,"pageType":17,"contentSource":24,"contentRef":105,"sort":15,"status":16,"icon":53,"description":106,"httpMethod":64,"hasToc":27},"01a06b6a-057b-73aa-b844-b72a2faaca84","grp-api-claude-opus-4-6","llm\u002Fclaude-opus-4-6","Claude Opus 4.6 Economy","01a06b6b-616f-787b-a23f-743e00db0302","Claude Opus 4.6 is Anthropic’s most powerful officially released Opus model, designed for complex reasoning, long-horizon agentic coding, and highly autonomous professional workflows. It supports text, image, and file inputs and produces text output. It features a 1-million-token context window, up to 128,000 output tokens, and built-in capabilities for adaptive thinking, tool use, and structured outputs. The model excels at advanced coding, browser and computer-use agents, enterprise knowledge work, financial and legal analysis, and multi-step tasks requiring sustained judgment and high reliability.",{"id":108,"groupId":109,"locale":4,"slug":110,"title":111,"pageType":17,"contentSource":24,"contentRef":112,"sort":15,"status":16,"icon":53,"description":113,"httpMethod":64,"hasToc":27},"01a06b6a-02eb-7415-82a8-8706f910d568","grp-api-claude-haiku-4-5","llm\u002Fclaude-haiku-4-5","Claude Haiku 4.5 Economy","01a06b6b-5195-7234-b601-490089fcf40a","Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence with significantly lower cost and latency than larger Claude models. With performance comparable to Claude Sonnet 4 across reasoning, coding, and computer-use tasks, it brings advanced capabilities to real-time, high-volume applications.\n\nAs the first Haiku model to support extended thinking, Haiku 4.5 offers adjustable reasoning depth, summarized or interleaved thinking, and tool-assisted workflows spanning coding, Bash, web search, and computer use. Scoring over 73% on SWE-bench Verified, it ranks among the world’s leading coding models while remaining highly responsive for sub-agent orchestration, parallel execution, and large-scale deployment.",{"id":115,"groupId":116,"locale":4,"slug":10,"title":117,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":118,"description":119,"hasToc":55,"children":120},"01a06b69-b126-7963-bca6-ca399c0bab4d","grp-section-chatgpt","ChatGPT","gpt-image","Gain production-grade API access to OpenAI’s flagship GPT-5.6 model suite. Moving beyond a one-size-fits-all design, GPT-5.6 introduces a clear tiered architecture featuring Sol (Flagship Reasoning), Terra (Balanced Workhorse), and Luna (Lightweight Speedster). GPT-5.6 Sol introduces breakthrough capabilities like Max Reasoning effort and Parallel Subagents, enabling the model to break down and execute ultra-complex coding, cybersecurity tasks, and deep analytical workflows in parallel.",[121,129,136,144],{"id":122,"groupId":123,"locale":4,"slug":124,"title":125,"pageType":17,"contentSource":24,"contentRef":126,"sort":15,"status":16,"icon":127,"description":128,"httpMethod":64,"hasToc":27},"01a079c7-905b-79cf-9d74-de2ba4984cac","grp-api-gpt-6-astra","llm\u002Fgpt-6-astra","GPT 6 Astra Economy","01a079c4-3dfe-72bb-8b17-b77f37a5675e","openai","GPT-6 Astra is OpenAI's flagship frontier model engineered for long-horizon, complex end-to-end enterprise workflows. Representing a major generational leap toward agentic intelligence, Astra integrates deep multi-step reasoning, advanced software engineering, and native computer-use capabilities to navigate software interfaces and execute multi-application workflows directly via GUI. Powered by symbolic world modeling and dynamic course-correction, it can break open-ended goals into actionable steps, handle ambiguous edge cases, and execute complex operations across unstructured environments. Setting new benchmarks on ARC-AGI-3 as well as complex financial, scientific, and coding tasks, GPT-6 Astra is built for enterprise agent orchestration, autonomous software development, financial intelligence, and complex legal\u002Fbusiness decision support.",{"id":130,"groupId":131,"locale":4,"slug":132,"title":133,"pageType":17,"contentSource":24,"contentRef":134,"sort":15,"status":16,"icon":127,"description":135,"httpMethod":64,"hasToc":27},"01a06b6a-0721-766d-8f5e-82fb5787bf77","grp-api-gpt-5-6-terra","llm\u002Fgpt-5.6-terra","GPT 5.6 Terra Economy","01a06b6b-6ba2-7895-9219-18aadef3fa83","GPT-5.6 Terra is the balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is well suited for everyday coding, reasoning, and agentic workflows, offering a strong balance of quality, latency, and cost for general production use.",{"id":137,"groupId":138,"locale":4,"slug":139,"title":140,"pageType":17,"contentSource":24,"contentRef":141,"sort":15,"status":16,"icon":142,"description":143,"httpMethod":64,"hasToc":27},"01a06b6a-07f4-7f04-950f-293316fd7b97","grp-api-gpt-5-6-sol","llm\u002Fgpt-5.6-sol","GPT 5.6 Sol Economy","01a06b6b-70c2-755b-b466-4cf804787a0d","chatpgt","GPT-5.6 Sol is the flagship model in OpenAI’s GPT-5.6 series. It is specifically designed for complex reasoning, coding, and agentic workflows, and particularly excels at multi-step problem solving, command-line assistance, and high-quality software tasks. This model is recommended when output quality and reliability matter more than raw throughput.",{"id":145,"groupId":146,"locale":4,"slug":147,"title":148,"pageType":17,"contentSource":24,"contentRef":149,"sort":15,"status":16,"icon":127,"description":150,"httpMethod":64,"hasToc":27},"01a06b6a-0d5c-7f9b-8568-d37c27e6bd3d","grp-api-gpt-5-5","llm\u002Fgpt-5.5","GPT 5.5 Economy","01a06b6b-9174-74cf-b3f7-d3ebc4ea8184","GPT-5.5 is a frontier model released by OpenAI on April 23, 2026. It features a context window of over 1 million tokens—922,000 input tokens and 128,000 output tokens—and supports both text and image inputs. The model scores 88.7% on SWE-bench Verified and 92.4% on MMLU, while reducing hallucinations by 60% compared with GPT-5.4. It excels in agentic coding, computer use, and deep research, while maintaining per-token latency comparable to GPT-5.4.",{"id":152,"groupId":153,"locale":4,"slug":10,"title":154,"pageType":32,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":155,"hasToc":27,"children":156},"01a08411-7d14-7652-b46f-f36ca5efdb4a","grp-section-gemini","Gimini","gemini",[157,164,171],{"id":158,"groupId":159,"locale":4,"slug":160,"title":161,"pageType":17,"contentSource":24,"contentRef":162,"sort":15,"status":16,"icon":155,"description":163,"httpMethod":64,"hasToc":27},"01a0842c-35e4-7650-b70f-4ae486d5f205","grp-api-gemini-3-7-flash","llm\u002Fgemini-3.7-flash","Gemini 3.7 Flash","01a08017-4fe3-7517-9b83-8592c565e598","Gemini 3.7 Flash is Google's flagship high-efficiency workhorse model engineered for autonomous agentic workflows and native multimodal reasoning. Built on enhanced core algorithmic foundations, it delivers ultra-low latency inference alongside robust multi-step tool execution, adaptive task planning, and end-to-end software engineering. Operating at $0.75 per million input tokens, Gemini 3.7 Flash powers enterprise agent orchestration, low-latency conversational interactive applications, real-time code generation, and high-volume data processing pipelines.",{"id":165,"groupId":166,"locale":4,"slug":167,"title":168,"pageType":17,"contentSource":24,"contentRef":169,"sort":15,"status":16,"icon":155,"description":170,"httpMethod":64,"hasToc":27},"01a0842e-667e-78a2-a943-8ee9b7b84fc7","grp-api-gemini-3-6-flash","llm\u002Fgemini-3.6-flash","Gemini 3.6 Flash","01a06b6a-c7e8-76c4-b2fc-689005fccf0b","Gemini 3.6 Flash is Google's next-generation lightweight workhorse model released in July 2026. Built for agentic workflows, complex coding, and multimodal tasks, it supports a 1M token input context window and a 64K token output limit. Compared to 3.5 Flash, it reduces output token consumption by ~17% with streamlined reasoning steps and tool calls, significantly lowering overall cost and latency for agent execution. Natively handling text, image, video, audio, and PDF inputs, it excels at computer use and multi-tool orchestration.",{"id":172,"groupId":173,"locale":4,"slug":174,"title":175,"pageType":17,"contentSource":24,"contentRef":176,"sort":15,"status":16,"icon":155,"description":177,"httpMethod":64,"hasToc":27},"01a0842f-0213-7969-9788-b4f50b032dc1","grp-api-gemini-3-5-flash","llm\u002Fgemini-3.5-flash","Gemini 3.5 Flash","01a06b6a-ad1f-7d0c-86ba-438a827b41b5","Gemini 3.5 Flash is Google's high-efficiency, lightweight multimodal workhorse model. Engineered for high-throughput, low-latency agentic workflows, code generation, and multimodal understanding, it features a 1M token context window. Natively processing text, image, video, audio, and document inputs, it delivers exceptional inference speed and cost-efficiency alongside strong tool-use and multilingual capabilities—ideal for enterprise API integrations and real-time interactive applications.",{"id":179,"groupId":180,"locale":4,"slug":10,"title":181,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":182,"description":183,"hasToc":55,"children":184},"01a06b69-aaf1-7ec6-be48-5ca932965a4c","grp-section-glm","GLM","zai","Gain production-grade API access to Z.ai’s (Zhipu AI) flagship foundation model, GLM-5.3. Leveraging the efficient MoE architecture combined with massive post-training scaling across high-complexity, long-horizon environments, GLM-5.3 delivers a massive leap in reasoning intelligence and compute efficiency. It demonstrates frontier capabilities in autonomous software engineering, complex code generation, terminal execution (Terminal-Bench), and defensive cybersecurity threat analysis.",[185,192,199],{"id":186,"groupId":187,"locale":4,"slug":188,"title":189,"pageType":17,"contentSource":24,"contentRef":190,"sort":15,"status":16,"icon":182,"description":191,"httpMethod":64,"hasToc":27},"01a06b69-eae5-712a-a5c7-8c723f7d1145","grp-api-glm-5-3-flash","llm\u002Fglm-5.3-flash","GLM 5.3 Flash","01a06b6a-b71d-73f7-8ab2-c3c64a80abf7","GLM 5.3 Flash is Zhipu AI's next-generation high-speed, lightweight workhorse model. Engineered for high-throughput, low-latency agentic workflows, code generation, and multimodal tasks, it features native long context support with significantly enhanced inference throughput and extreme cost-efficiency. It excels in function calling, instruction following, logical reasoning, and multilingual understanding—ideal for enterprise API integrations, real-time interactive apps, and automated workflows.",{"id":193,"groupId":194,"locale":4,"slug":195,"title":196,"pageType":17,"contentSource":24,"contentRef":197,"sort":15,"status":16,"icon":182,"description":198,"httpMethod":64,"hasToc":27},"01a06b69-f231-711b-a019-57669562e441","grp-api-glm-5-3","llm\u002Fglm-5.3","GLM 5.3","01a06b6a-e2bf-7f47-97d8-4e1578c645e5","GLM 5.3 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering, and complex multi-step automation.\n\nReasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is particularly strong at coding and tool use across long-running tasks, able to maintain engineering context and follow standards consistently through a full development workflow, from requirements to multi-platform deployment, in a single task.",{"id":200,"groupId":201,"locale":4,"slug":202,"title":203,"pageType":17,"contentSource":24,"contentRef":204,"sort":15,"status":16,"icon":182,"description":205,"httpMethod":64,"hasToc":27},"01a06b69-e83d-7f78-8532-621345a762b7","grp-api-glm-5-2","llm\u002Fglm-5.2","GLM 5.2","01a06b6a-a802-792b-893a-83c6bb5191d2","GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering, and complex multi-step automation.\n\nReasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is particularly strong at coding and tool use across long-running tasks, able to maintain engineering context and follow standards consistently through a full development workflow, from requirements to multi-platform deployment, in a single task.",{"id":207,"groupId":208,"locale":4,"slug":10,"title":209,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":210,"description":211,"hasToc":55,"children":212},"01a06b69-b04b-7e6e-9b72-7dc6f29a98ef","grp-section-deepseek","DeepSeek","deepseek","Gain production-grade API access to DeepSeek’s flagship DeepSeek V4 model family. Setting industry standards for open-weights performance, the DeepSeek V4 suite features DeepSeek-V4-Pro (Flagship Reasoning) and DeepSeek-V4-Flash (High-Throughput Speedster).",[213,220],{"id":214,"groupId":215,"locale":4,"slug":216,"title":217,"pageType":17,"contentSource":24,"contentRef":218,"sort":15,"status":16,"icon":210,"description":219,"httpMethod":64,"hasToc":27},"01a06b69-e760-73b9-b255-d1ea72fee4b3","grp-api-deepseek-v4-pro","llm\u002Fdeepseek-v4-pro","Deepseek V4 Pro 0813","01a06b6a-a2fb-7110-8278-bde106d3a6a7","DeepSeek-V4-Pro-0813 is a flagship MoE model released on August 13, 2026, serving as the official GA release built for complex logical reasoning, code construction, and AI Agent collaboration. Architecture & Performance: 1.6T parameters (49B active), built-in DSpark speculative decoding, with significantly enhanced inference throughput. Compatible with OpenAI\u002FAnthropic APIs, deeply integrated with DeepSeek Harness, and delivers outstanding performance on Terminal Bench.",{"id":221,"groupId":222,"locale":4,"slug":223,"title":224,"pageType":17,"contentSource":24,"contentRef":225,"sort":15,"status":16,"icon":210,"description":226,"httpMethod":64,"hasToc":27},"01a06b6a-0c88-7354-942c-77aa78ecf029","grp-api-deepseek-v4-flash","llm\u002Fdeepseek-v4-flash","Deepseek V4 Flash 0731","01a06b6b-8c88-77d5-a914-02d485613e30","DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts (MoE) model from DeepSeek, featuring 284B total parameters with 13B activated per token and a 1-million-token context window. Designed for fast inference and high-throughput workloads, it delivers strong reasoning and coding capabilities while maintaining excellent cost efficiency.\n\nIts hybrid attention architecture enables efficient long-context processing. The model supports **high** and **xhigh** reasoning levels, with **xhigh** representing the maximum reasoning effort. It is ideal for coding assistants, conversational systems, and agentic workflows where responsiveness, scalability, and cost efficiency are essential.",{"id":228,"groupId":229,"locale":4,"slug":10,"title":230,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":231,"description":232,"hasToc":55,"children":233},"01a06b69-b8ae-78fe-b5e3-444762883eaf","grp-section-qwen","Qwen","qwen","Gain production-grade API access to Alibaba Cloud’s complete Qwen 3.8 and Qwen3 foundation model families. Setting global performance and cost-efficiency benchmarks, the Qwen suite features Qwen3.8-Max (Flagship 2.4T Reasoning & Coding Engine), Qwen3.7-Plus (Balanced Enterprise Workhorse), Qwen3.8-27B (Low-Latency Speedster), and Qwen-Image-3.0 (Photorealistic Visual Generator).",[234,241,248,255],{"id":235,"groupId":236,"locale":4,"slug":237,"title":238,"pageType":17,"contentSource":24,"contentRef":239,"sort":15,"status":16,"icon":231,"description":240,"httpMethod":64,"hasToc":27},"01a06b69-ec16-7054-a65c-b70d1db7dfef","grp-api-qwen3-8-max","llm\u002Fqwen3.8-max","Qwen 3.8 Max","01a06b6a-bc33-749a-a00d-e6563ac9f874","Qwen3.8-Max is Alibaba’s flagship model in the Qwen3.8 series, designed for agentic, text-based workflows. It excels at coding, debugging, office automation, productivity tasks, tool use, and long-horizon autonomous execution. With a 1-million-token context window and support for outputs of up to 64K tokens, it is ideal for processing large documents, repository-scale coding, multi-step planning, structured content generation, and complex workflows requiring sustained reasoning across hundreds or even thousands of steps.",{"id":242,"groupId":243,"locale":4,"slug":244,"title":245,"pageType":17,"contentSource":24,"contentRef":246,"sort":15,"status":16,"icon":231,"description":247,"httpMethod":64,"hasToc":27},"01a06b69-f305-7365-a340-b25d236b14c0","grp-api-qwen3-7-max","llm\u002Fqwen3.7-max","Qwen 3.7 Max","01a06b6a-e7f3-73a2-be3d-36a7334c9cde","Qwen3.7-Max is Alibaba’s flagship model in the Qwen3.7 series, designed for agentic, text-based workflows. It excels at coding, debugging, office automation, productivity tasks, tool use, and long-horizon autonomous execution. With a 1-million-token context window and support for outputs of up to 64K tokens, it is ideal for processing large documents, repository-scale coding, multi-step planning, structured content generation, and complex workflows requiring sustained reasoning across hundreds or even thousands of steps.",{"id":249,"groupId":250,"locale":4,"slug":251,"title":252,"pageType":17,"contentSource":24,"contentRef":253,"sort":15,"status":16,"icon":231,"description":254,"httpMethod":64,"hasToc":27},"01a06b69-f3dc-7ef2-aa67-fd820218cd07","grp-api-qwen3-7-plus","llm\u002Fqwen3.7-plus","Qwen 3.7 Plus","01a06b6a-ed15-70a6-aee8-7a58aae503a0","Qwen3.7-Plus is a cost-effective model in Alibaba’s Qwen3.7 series, supporting text and image inputs with text output. It combines enhanced vision-language capabilities with full-stack, agentic intelligence for coding, tool use, and productivity workflows. Its key strength is multimodal, interactive agent functionality: it can understand real-world scenes, interpret on-screen content, interact with graphical interfaces, generate code from visual references, and autonomously navigate mobile apps from end to end.",{"id":256,"groupId":257,"locale":4,"slug":258,"title":259,"pageType":17,"contentSource":24,"contentRef":260,"sort":15,"status":16,"icon":231,"description":261,"httpMethod":64,"hasToc":27},"01a06b69-ea10-7cb2-af7a-8da1657886c8","grp-api-qwen3-5-omni-plus","llm\u002Fqwen3.5-omni-plus","Qwen 3.5 Omni Plus","01a06b6a-b221-7009-bf1c-2558df585cc6","Qwen 3.5 Omni Plus is Alibaba Cloud's next-generation native multimodal (Omni) flagship model. Powered by an end-to-end unified architecture, it supports real-time full-duplex interaction across text, image, audio, and video inputs and outputs. It features ultra-low latency streaming inference, expressive speech synthesis, joint vision-audio reasoning, and strong tool-use capability—ideal for real-time voice assistants, AI customer service, video interaction, and digital human agents.",{"id":263,"groupId":264,"locale":4,"slug":10,"title":265,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":266,"description":267,"hasToc":55,"children":268},"01a06b69-b1ff-78f5-99ab-b99371276f88","grp-section-minimax","MiniMax","minimax","API access to MiniMax M3 & M2.7 is live. M3 features MSA attention with a 1M context window, native multimodality, and elite coding. M2.7 excels at self-evolving Agent Teams and complex office workflows. Powered by low latency and Prompt Caching for studio-ready enterprise pipelines.",[269,276],{"id":270,"groupId":271,"locale":4,"slug":272,"title":273,"pageType":17,"contentSource":24,"contentRef":274,"sort":15,"status":16,"icon":266,"description":275,"httpMethod":64,"hasToc":27},"01a06b69-ef86-711a-8e3f-2e19c7cf0ab9","grp-api-MiniMax-M3","llm\u002FMiniMax-M3","MiniMax M3","01a06b6a-d21a-7c17-b886-de7bd6ba9060","MiniMax-M3 is MiniMax’s latest M-series multimodal foundation model, built for agentic reasoning, tool use, coding, and long-context tasks. It supports text, image, and video inputs with text output, offering a 1M-token context window, extended thinking, function calling, and structured outputs. With strong capabilities in long-horizon agent workflows, software development, multimodal understanding, and extended response generation, MiniMax-M3 is ideal for autonomous agents, coding assistants, document and video analysis, and production-grade applications that require massive context at a competitive cost.",{"id":277,"groupId":278,"locale":4,"slug":279,"title":280,"pageType":17,"contentSource":24,"contentRef":281,"sort":15,"status":16,"icon":266,"description":282,"httpMethod":64,"hasToc":27},"01a06b69-f07e-7ea1-b3da-c778997970fe","grp-api-MiniMax-M2-7","llm\u002FMiniMax-M2.7","MiniMax M2.7","01a06b6a-d824-764c-ac7f-ce3be1bff924","MiniMax-M2.7 is a next-generation large language model built for autonomous real-world productivity and continuous improvement. Through advanced agentic capabilities and multi-agent collaboration, it can plan, execute, evaluate, and refine complex tasks in dynamic environments while actively contributing to its own evolution.\n\nOptimized for production-grade workflows, M2.7 excels at live debugging, root-cause analysis, financial modeling, and end-to-end document creation across Word, Excel, and PowerPoint. It achieves 56.2% on SWE-Pro, 57.0% on Terminal Bench 2, and a 1495 ELO rating on GDPval-AA, establishing a new benchmark for multi-agent systems operating in real-world digital workflows.",{"id":284,"groupId":285,"locale":4,"slug":10,"title":286,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":287,"description":288,"hasToc":55,"children":289},"01a06b69-b2e3-70ff-847a-e01b01d3ffe0","grp-section-kimi","Kimi","kimi","Gain production-grade API access to Moonshot AI’s flagship Kimi K3 and K2 model series. Built to process massive context windows with zero loss alongside advanced Context Caching for lower latency and cost, Kimi excels at complex code generation, multi-step problem solving, automated research, and dynamic tool orchestration.",[290],{"id":291,"groupId":292,"locale":4,"slug":293,"title":294,"pageType":17,"contentSource":24,"contentRef":295,"sort":15,"status":16,"icon":287,"description":296,"httpMethod":64,"hasToc":27},"01a06b69-f157-768a-90e2-f8b1ed5340e9","grp-api-kimi-k3","llm\u002Fkimi-k3","Kimi K3","01a06b6a-dd6e-77fe-ab72-4f93f9941d80","Kimi K3 is Kimi’s most capable flagship model to date. With 2.8 trillion parameters, it is built on the Kimi Delta Attention (KDA) hybrid linear attention architecture and Attention Residuals technology. It natively supports visual understanding and features a 1-million-token context window.\n\nAs the world’s first open-source model at the 3-trillion-parameter scale, Kimi K3 is designed for frontier AI use cases, including long-horizon coding, knowledge work, and reasoning.",{"id":298,"groupId":299,"locale":4,"slug":10,"title":300,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":301,"description":302,"hasToc":55,"children":303},"01a06b69-a774-7e34-b542-4589456c4d68","grp-section-hy","Tencent Hy","hunyuan","Gain production-grade API access to Tencent’s flagship Hunyuan 3.0 (Hy3) foundation language model. Featuring an innovative hybrid fast-and-slow thinking process and supporting up to a 256K context window, Hy3 achieves breakthrough execution across autonomous agent workflows, complex code synthesis (SWE-Bench), deep reasoning, and enterprise workspace automation.",[304,311],{"id":305,"groupId":306,"locale":4,"slug":307,"title":308,"pageType":17,"contentSource":24,"contentRef":309,"sort":15,"status":16,"icon":301,"description":310,"httpMethod":64,"hasToc":27},"01a06b69-eeaf-7923-a9b3-747c22cb2d0a","grp-api-hy4-preview","llm\u002Fhy4-preview","Tencent Hy4 Preview","01a06b6a-ccfd-7547-9e42-99ccbf61f6df","Tencent Hy4 preview (Tencent Hunyuan 4 preview) is Tencent's next-generation 770B MoE open-source flagship model. Comprising 770 billion total parameters and 49 billion active parameters, it supports a 1-million-token (1M) context window. Architectural innovations include Gated DeepSeek Sparse Attention (Gated DSA) with IndexCache, identity Hyper-Connections (iHC), and a built-in 10B MTP layer for speculative decoding.",{"id":312,"groupId":313,"locale":4,"slug":314,"title":315,"pageType":17,"contentSource":24,"contentRef":316,"sort":15,"status":16,"icon":301,"description":317,"httpMethod":64,"hasToc":27},"01a06b69-ece9-7ca4-9aaa-2d9ca3aa52cf","grp-api-hy3","llm\u002Fhy3","Tencent Hy3","01a06b6a-c2bf-70c0-951d-e824259a0489","Tencent Hy3 (Tencent Hunyuan 3) is Tencent's next-generation flagship MoE model. Built on a 295B-parameter MoE architecture with 21B active parameters and Multi-Token Prediction (MTP), it is engineered for agentic workflows, complex coding, long-context processing, and deep reasoning.",{"id":319,"groupId":320,"locale":4,"slug":10,"title":321,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":322,"description":323,"hasToc":55,"children":324},"01a06b69-a852-794d-bb81-f7f809f6325f","grp-section-mimo-v2.5","MiMo V2.5","xiaomimimo","Xiaomi MiMo is Xiaomi's next-generation AI model family, serving as the core intelligence foundation for its \"Human x Car x Home\" ecosystem. Built for native multimodal understanding (text, vision, audio, and video), complex logical reasoning, code generation, and agentic workflows, it supports up to a 1-million-token context window with ultra-fast streaming inference.",[325],{"id":326,"groupId":327,"locale":4,"slug":328,"title":329,"pageType":17,"contentSource":24,"contentRef":330,"sort":15,"status":16,"icon":322,"description":331,"httpMethod":64,"hasToc":27},"01a08015-aa41-773f-9ddb-af655a415914","grp-api-xiaomi-mimo-v2-5-pro","llm\u002Fxiaomi\u002Fmimo-v2.5-pro","MiMo V2.5 Pro","01a08017-5592-7ce3-bf75-809f549a1e15","MiMo V2.5 Pro (Xiaomi MiMo V2.5 Pro) is Xiaomi's flagship open-source Mixture-of-Experts (MoE) language model engineered for advanced agentic workloads. Featuring 1.02 trillion total parameters and 42 billion active parameters, it supports a 1-million-token (1M) context window. Its architecture integrates a hybrid attention mechanism (interleaved Sliding Window and Global Attention) alongside a native Multi-Token Prediction (MTP) layer, achieving frontier-tier intelligence with extreme token efficiency.",{"id":333,"groupId":334,"locale":4,"slug":10,"title":335,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"topTab":17,"children":336},"01a06b0c-7859-78ae-9677-d20d9fc7b712","grp-section-image-models","Image Models",[337,362,380,401,413,433,452],{"id":338,"groupId":339,"locale":4,"slug":10,"title":340,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":127,"description":341,"hasToc":55,"children":342},"01a08514-16f4-7c81-9229-3f1b6ddd692a","grp-section-gpt-image-2.5","GPT Images 2.5","GPT Images 2.5 is OpenAI's flagship commercial image generation and precise editing model available in both Standard (optimized for photographic and general styles) and Ultra (engineered for crisp text rendering and extreme definition) editions. Supporting up to 10 reference images, it ensures strict multi-image identity, material, and logo consistency across dynamic compositions.",[343,350,356],{"id":344,"groupId":345,"locale":4,"slug":346,"title":347,"pageType":17,"contentSource":24,"contentRef":348,"sort":15,"status":16,"icon":118,"description":349,"httpMethod":64,"hasToc":27},"01a08514-2ae7-7044-80ab-cb34c3b0f225","grp-api-openai-gpt-image-2-5-economy","image\u002Fopenai\u002Fgpt-image-2-5-economy","GPT Image 2.5 Economy","01a08512-9d90-707e-8467-a1d722ce08d0","GPT Images 2.5 is OpenAI's flagship commercial image generation and precise editing model available in both Standard (optimized for photographic and general styles) and Ultra (engineered for crisp text rendering and extreme definition) editions. Supporting up to 10 reference images, it ensures strict multi-image identity, material, and logo consistency across dynamic compositions. With built-in Sketch-to-Image control and smart Templates, users can convert hand-drawn layouts into polished designs and automate e-commerce product replacement. Delivering native 4K output and precise typography control, GPT Images 2.5 powers cross-border e-commerce visual pipelines, brand marketing assets, UI\u002FUX concept design, and high-converting commercial media.",{"id":351,"groupId":352,"locale":4,"slug":353,"title":354,"pageType":17,"contentSource":24,"contentRef":355,"sort":15,"status":16,"icon":118,"httpMethod":64,"hasToc":27},"01a09017-0042-7fdf-8c24-58d41f6a38ea","grp-api-openai-gpt-image-2-5-flare-economy","image\u002Fopenai\u002Fgpt-image-2-5-flare-economy","GPT Image 2.5 Flare Economy","01a09015-98ae-7cad-bc77-b17834c0e303",{"id":357,"groupId":358,"locale":4,"slug":359,"title":360,"pageType":17,"contentSource":24,"contentRef":361,"sort":15,"status":16,"icon":118,"httpMethod":64,"hasToc":27},"01a09017-01ef-7688-a3ad-84c5025f3204","grp-api-openai-gpt-image-2-5-sunburst-economy","image\u002Fopenai\u002Fgpt-image-2-5-sunburst-economy","GPT Image 2.5 Sunburst Economy","01a09015-a2aa-7e86-88ea-a2d19f2fef80",{"id":363,"groupId":364,"locale":4,"slug":10,"title":365,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":118,"description":366,"hasToc":55,"children":367},"01a06b69-ae94-7a33-a93f-1dcf863fd402","grp-section-gpt-image-2","GPT Image 2","GPT Image 2 introduces native visual reasoning (Thinking Mode). Delivering up to 2K crisp resolution, it excels in high-fidelity photorealism, multi-language typography, and structured UI\u002Finfographic generation. Furthermore, GPT Image 2 delivers industry-leading character\u002Fstyle consistency across multi-image sets and supports precise instruction-guided editing without destroying underlying compositions.",[368,374],{"id":369,"groupId":370,"locale":4,"slug":371,"title":365,"pageType":17,"contentSource":24,"contentRef":372,"sort":15,"status":16,"icon":127,"description":373,"httpMethod":64,"hasToc":27},"01a06b69-d695-7bfa-8fb7-8b1b0470bcef","grp-api-openai-gpt-image-2","image\u002Fopenai\u002Fgpt-image-2","01a06b6a-4b51-7979-a729-1f712b70ae07","OpenAI's GPT Image 2 raw image model can generate high-quality images based on natural language prompts. It provides a ready-to-use REST inference API, offering excellent performance, no cold start, and affordability.",{"id":375,"groupId":376,"locale":4,"slug":377,"title":378,"pageType":17,"contentSource":24,"contentRef":379,"sort":15,"status":16,"icon":127,"description":373,"httpMethod":64,"hasToc":27},"01a06b69-d76c-7f77-87c1-9e2594a639aa","grp-api-openai-gpt-image-2-text-to-image","image\u002Fopenai\u002Fgpt-image-2\u002Ftext-to-image","GPT Image 2 Text-to-Image","01a06b6a-5067-7246-ad37-59f693b48db3",{"id":381,"groupId":382,"locale":4,"slug":10,"title":383,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":384,"description":385,"hasToc":55,"children":386},"01a06b69-a69b-7303-ac46-001594b17cac","grp-section-seedream5.0pro","Seedream 5.0 Pro","bytedance","Gain production-grade API access to ByteDance’s flagship image foundation model, Seedream 5.0 Pro. Engineered as an industrial benchmark for visual synthesis, 5.0 Pro delivers photorealistic lighting, accurate material textures, and precise composition control—natively rendering up to 4K commercial resolution. The model introduces exact HEX color code alignment to enforce strict brand compliance, combined with state-of-the-art multi-language typography and robust multi-reference identity preservation.",[387,394],{"id":388,"groupId":389,"locale":4,"slug":390,"title":391,"pageType":17,"contentSource":24,"contentRef":392,"sort":15,"status":16,"icon":384,"description":393,"httpMethod":64,"hasToc":27},"01a06b69-df50-72cf-950b-0d6dca154420","grp-api-seedream-5-0-pro-text-to-image","image\u002Fseedream-5.0-pro\u002Ftext-to-image","Seedream 5.0 Pro Text-to-Image Spicy","01a06b6a-7a26-7d2d-992b-29f301770131","Seedream 5.0 Pro Text to Image Spicy (High Saturation\u002FDynamic Range Edition) is a uncensored and professional-grade AI text-to-image model variant by Seedream AI, dedicated to extreme visual impact. Building upon the professional detail and control of the 5.0 Pro series, the Spicy edition is fine-tuned for maximized color saturation, dynamic range (HDR), dramatic lighting contrast, and powerful visual tension.",{"id":395,"groupId":396,"locale":4,"slug":397,"title":398,"pageType":17,"contentSource":24,"contentRef":399,"sort":15,"status":16,"icon":384,"description":400,"httpMethod":64,"hasToc":27},"01a06b69-e114-7c33-b8f3-bbbdb09c8fb7","grp-api-seedream-5-0-pro-edit","image\u002Fseedream-5.0-pro\u002Fedit","Seedream 5.0 Pro Edit Spicy","01a06b6a-847c-7c25-8676-0ee0b0b70a21","Seedream 5.0 Pro Edit Spicy is a uncensored and professional-grade AI image edit model variant by Seedream AI, dedicated to extreme visual impact. Building upon the professional detail and control of the 5.0 Pro series, the Spicy edition is fine-tuned for maximized color saturation, dynamic range (HDR), dramatic lighting contrast, and powerful visual tension.",{"id":402,"groupId":403,"locale":4,"slug":10,"title":404,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":384,"description":405,"hasToc":55,"children":406},"01a06b69-abd2-7228-8253-c56246f2ccef","grp-section-seedream5.0lite","Seedream 5.0 Lite","Seedream 5.0 Lite brings ByteDance’s cutting-edge image generation capabilities to high-velocity, cost-sensitive production pipelines. Built as a lightweight, high-throughput model in the Seedream 5.0 series, Lite delivers accelerated rendering speeds and a significantly lower cost per image while retaining the essential strengths of the flagship architecture—such as sharp multi-language typography, structured poster layout, and solid compositional fidelity.",[407],{"id":408,"groupId":409,"locale":4,"slug":410,"title":404,"pageType":17,"contentSource":24,"contentRef":411,"sort":15,"status":16,"icon":384,"description":412,"httpMethod":64,"hasToc":27},"01a06b69-de7a-748c-ae62-d7da1f275f78","grp-api-bytedance-seedream-5-0","image\u002Fbytedance\u002Fseedream-5-0","01a06b6a-750f-73a2-a240-94cad7a4e907","Seedream 5.0 by is a state-of-the-art text-to-image model with enhanced typography, clear text rendering for posters and brand visuals, superior prompt adherence, and up to 4K resolution. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.",{"id":414,"groupId":415,"locale":4,"slug":10,"title":416,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":417,"description":418,"hasToc":55,"children":419},"01a06b69-af70-7ae4-84e5-82fe87cb6e2c","grp-section-nano-banana","Nano Banana Pro","nanobanana","Designed for real-world creative and enterprise workflows, Nano Banana redefines visual content creation through native multimodal editing and real-time world knowledge integration. The model suite supports crisp 1K, 2K, and direct native 4K (up to 4096×2304) resolutions without external upscaling, excelling at intricate multilingual text rendering, studio layouts, and brand graphics.",[420,426],{"id":421,"groupId":422,"locale":4,"slug":423,"title":416,"pageType":17,"contentSource":24,"contentRef":424,"sort":15,"status":16,"icon":417,"description":425,"httpMethod":64,"hasToc":27},"01a06b69-f736-7313-a619-524d4b7605aa","grp-api-google-gemini-3-pro-image","image\u002Fgoogle\u002Fgemini-3-pro-image","01a06b6b-0217-7500-ba80-521b864a7678","Google Nano Banana Pro (Gemini 3.0 Pro Image) supports image editing and can output 4K resolution results. It offers a ready-to-use REST inference API, excellent performance, no cold start, and affordable price.",{"id":427,"groupId":428,"locale":4,"slug":429,"title":430,"pageType":17,"contentSource":24,"contentRef":431,"sort":15,"status":16,"icon":417,"description":432,"httpMethod":64,"hasToc":27},"01a06b69-f80c-7009-ade3-05bef2d7e4bf","grp-api-google-gemini-3-pro-image-text-to-image","image\u002Fgoogle\u002Fgemini-3-pro-image\u002Ftext-to-image","Nano Banana Pro Text-to-Image","01a06b6b-0739-7b27-a019-423095638bb0","Google Nano Banana Pro (Gemini 3.0 Pro Image) supports text-to-image and can output 4K resolution results. It offers a ready-to-use REST inference API, excellent performance, no cold start, and affordable price.",{"id":434,"groupId":435,"locale":4,"slug":10,"title":436,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":417,"description":437,"hasToc":55,"children":438},"01a06b69-b99d-7354-81f2-1aae0ace956c","grp-section-nano-banana-2","Nano Banana 2","Gain production-grade API access to Nano Banana 2 (gemini-3.1-flash-image), powered by Google’s latest Gemini 3.1 Flash Image foundation model. Redefining high-throughput visual creation, Nano Banana 2 combines Pro-tier visual fidelity with ultra-fast Flash execution speed, supporting direct native 4K resolution (up to 4096×2304) without external upscaling.",[439,446],{"id":440,"groupId":441,"locale":4,"slug":442,"title":443,"pageType":17,"contentSource":24,"contentRef":444,"sort":15,"status":16,"icon":417,"description":445,"httpMethod":64,"hasToc":27},"01a06b69-f657-7df5-8e00-bc2d799f638f","grp-api-google-gemini-3-1-flash-image-text-to-image","image\u002Fgoogle\u002Fgemini-3-1-flash-image\u002Ftext-to-image","Nano Banana 2 Text-to-Image","01a06b6a-fcf5-7a7e-aba8-345e35fa2c7e","Nano Banana 2 Text-to-Image (Gemini 3.1 Flash Image) is Google’s next-generation AI model for image editing and generation, making visual creation as simple and intuitive as describing it in words. Built on Google’s cutting-edge computer vision and generative AI technologies, it combines precise control, creative flexibility, and deep semantic understanding to deliver professional-grade image editing and generation.",{"id":447,"groupId":448,"locale":4,"slug":449,"title":436,"pageType":17,"contentSource":24,"contentRef":450,"sort":15,"status":16,"icon":417,"description":451,"httpMethod":64,"hasToc":27},"01a06b69-fa91-74e4-94ee-d1d445ad723d","grp-api-google-gemini-3-1-flash-image","image\u002Fgoogle\u002Fgemini-3-1-flash-image","01a06b6b-1956-74bc-9c40-e855530a3bca","Nano Banana 2 (Gemini 3.1 Flash Image) is Google’s next-generation AI model for image editing and generation, making visual creation as simple and intuitive as describing it in words. Built on Google’s cutting-edge computer vision and generative AI technologies, it combines precise control, creative flexibility, and deep semantic understanding to deliver professional-grade image editing and generation.",{"id":453,"groupId":454,"locale":4,"slug":10,"title":455,"pageType":32,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":231,"description":232,"hasToc":27,"children":456},"01a0ad8d-620e-766c-9e31-f1537bf662b4","grp-section-qwen-image","Qwen Image",[457,464],{"id":458,"groupId":459,"locale":4,"slug":460,"title":461,"pageType":17,"contentSource":24,"contentRef":462,"sort":15,"status":16,"icon":231,"description":463,"httpMethod":64,"hasToc":27},"01a079c7-936d-7f9a-b882-095b3e6ffe54","grp-api-aliyun-qwen-image-3-0","llm\u002Faliyun\u002Fqwen-image-3-0","Qwen Image 3.0","01a06c19-1939-732f-b053-71ddc2534e90","Qwen Image 3.0 is Alibaba's next-generation open-weights AI image generation and editing model built on a Diffusion Transformer (DiT) architecture. Supporting both text-to-image (T2I) and image-to-image instruction editing (I2I), it excels at rich content rendering, hyper-realistic detail, and deep contextual understanding. It achieves industry-leading text-in-image typography, rendering clear small text down to 10px across 12 native languages and 20+ fonts for complex multi-column layouts, UI designs, and graphic documents. With open weights available for self-hosting, Qwen Image 3.0 provides highly cost-effective, high-precision visual generation for localized advertising, product UI prototyping, e-commerce graphics, and creative publishing.",{"id":465,"groupId":466,"locale":4,"slug":467,"title":468,"pageType":17,"contentSource":24,"contentRef":469,"sort":15,"status":16,"icon":231,"description":470,"httpMethod":64,"hasToc":27},"01a079c7-94f8-7197-abec-69542652e4df","grp-api-aliyun-qwen-image-3-0-pro","llm\u002Faliyun\u002Fqwen-image-3-0-pro","Qwen Image 3.0 Pro","01a06c19-1e83-7226-8933-91ccb77c3bd7","Qwen Image 3.0 Pro is Alibaba's flagship third-generation AI image generation model built on a Diffusion Transformer architecture. Engineered for high information density and professional production, it accepts ultra-long prompts of up to 4,500 tokens, allowing a single pass to generate complex layouts like newspaper front pages, multi-panel storyboards, academic papers, and detailed UI interfaces. It delivers micro-level detail and industry-leading typography, rendering legible small text down to 10px across 12 native languages and 20+ fonts.",{"id":472,"groupId":473,"locale":4,"slug":10,"title":474,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"topTab":17,"children":475},"01a06b0c-7933-7314-8cf6-b2b99a878202","grp-section-video-models","Video Models",[476,502,549,575,624,644,670,761,795],{"id":477,"groupId":478,"locale":4,"slug":10,"title":479,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":384,"description":480,"hasToc":55,"children":481},"01a06b69-acac-77bc-9499-8f67b13af306","grp-section-seedance2.5","Seedance 2.5","ByteDance’s next-generation video foundation model, Seedance 2.5, is now officially live. Built to push the boundaries of long-form storytelling, it produces up to 30 seconds of native, single-shot video in a single pass—guided by text prompts, a single image, or up to 50 multimodal reference inputs (images, video clips, and audio).",[482,488,495],{"id":483,"groupId":484,"locale":4,"slug":485,"title":479,"pageType":17,"contentSource":24,"contentRef":486,"sort":15,"status":16,"icon":384,"description":487,"httpMethod":64,"hasToc":27},"01a06b6a-09a4-7d42-ab4e-592e0731ee76","grp-api-bytedance-seedance-2-5","video\u002Fbytedance\u002Fseedance-2-5","01a06b6b-7d09-73ca-a966-62fa2fd2ea57","Seedance 2.5 is multimodal video generation from reference images, videos, and audio. Supports video editing and extension.",{"id":489,"groupId":490,"locale":4,"slug":491,"title":492,"pageType":17,"contentSource":24,"contentRef":493,"sort":15,"status":16,"icon":384,"description":494,"httpMethod":64,"hasToc":27},"01a06b6a-0a79-7de2-aef1-6ecb61f67f1e","grp-api-bytedance-seedance-2-5-text-to-video","video\u002Fbytedance\u002Fseedance-2-5\u002Ftext-to-video","Seedance 2.5 Text-to-Video","01a06b6b-823c-708e-b138-96025c4e6138","Seedance 2.5 Text-to-Video is a multimodal video generation tool based on reference text, supporting video editing and extension capabilities.",{"id":496,"groupId":497,"locale":4,"slug":498,"title":499,"pageType":17,"contentSource":24,"contentRef":500,"sort":15,"status":16,"icon":384,"description":501,"httpMethod":64,"hasToc":27},"01a06b6a-0bb0-75ca-906e-1cd21386920f","grp-api-bytedance-seedance-2-5-image-to-video","video\u002Fbytedance\u002Fseedance-2-5\u002Fimage-to-video","Seedance 2.5 Image-to-Video","01a06b6b-8762-7d33-b60a-23dcd06f659c","Seedance 2.5 Image-to-Video is a multimodal video generation tool based on reference images, supporting video editing and extension capabilities.",{"id":503,"groupId":504,"locale":4,"slug":10,"title":505,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":384,"description":506,"hasToc":55,"children":507},"01a06b69-a933-76ce-9048-8f921f3ee9de","grp-section-seedance2.0","Seedance 2.0","Gain production-grade API access to ByteDance’s landmark Seedance 2.0 video generation model. Built on a unified multimodal joint audio-video architecture, Seedance 2.0 supports seamless quad-modal prompting across text, images, video clips, and audio. Now upgraded with native 4K resolution rendering, Seedance 2.0 offers flexible, transparent pricing backed by enterprise-grade SLA reliability and instant API key provision—allowing production teams to build scalable video workflows with zero friction.",[508,515,522,529,536,543],{"id":509,"groupId":510,"locale":4,"slug":511,"title":512,"pageType":17,"contentSource":24,"contentRef":513,"sort":15,"status":16,"icon":384,"description":514,"httpMethod":64,"hasToc":27},"01a06b69-fd14-7d47-9f2f-5a370453e425","grp-api-bytedance-seedance-2-0-fast","video\u002Fbytedance\u002Fseedance-2-0-fast","Seedance 2.0 Fast","01a06b6b-2a41-796c-af4d-41a2e52f0691","Seedance 2.0-fast offers rapid video generation. It generates 4-15 second videos with text prompts, supports various aspect ratios, audio generation, and enhanced web search capabilities.",{"id":516,"groupId":517,"locale":4,"slug":518,"title":519,"pageType":17,"contentSource":24,"contentRef":520,"sort":15,"status":16,"icon":384,"description":521,"httpMethod":64,"hasToc":27},"01a06b69-fde9-770b-bc11-cd179e9a3bff","grp-api-bytedance-seedance-2-0-text-to-video","video\u002Fbytedance\u002Fseedance-2-0\u002Ftext-to-video","Seedance 2.0 Text-to-Video","01a06b6b-300e-76b9-85c5-68fe4ff28dd4","Seedance 2.0 Text-to-Video offers the highest image quality. It generates 4-15 second videos with text prompts and supports various aspect ratios, audio generation, and enhanced web search capabilities.",{"id":523,"groupId":524,"locale":4,"slug":525,"title":526,"pageType":17,"contentSource":24,"contentRef":527,"sort":15,"status":16,"icon":384,"description":528,"httpMethod":64,"hasToc":27},"01a06b69-ff97-76c2-94d4-b3e6d90503c7","grp-api-bytedance-seedance-2-0-image-to-video","video\u002Fbytedance\u002Fseedance-2-0\u002Fimage-to-video","Seedance 2.0 Image-to-Video","01a06b6b-3aed-7164-a609-dee9411f90f6","Seedance 2.0 Image-to-Video offers the highest image quality. It generates 4-15 second videos with text prompts and supports various aspect ratios, audio generation, and enhanced web search capabilities.",{"id":530,"groupId":531,"locale":4,"slug":532,"title":533,"pageType":17,"contentSource":24,"contentRef":534,"sort":15,"status":16,"icon":384,"description":535,"httpMethod":64,"hasToc":27},"01a06b6a-006f-7e5d-95d6-8f6ac28e571a","grp-api-bytedance-seedance-2-0-fast-text-to-video","video\u002Fbytedance\u002Fseedance-2-0-fast\u002Ftext-to-video","Seedance 2.0 Fast Text-to-Video","01a06b6b-405b-7f6f-950c-3fe5586ef73b","Seedance 2.0 Fast Text-to-Video offers the highest image quality. It generates 4-15 second videos with text prompts and supports various aspect ratios, audio generation, and enhanced web search capabilities.",{"id":537,"groupId":538,"locale":4,"slug":539,"title":540,"pageType":17,"contentSource":24,"contentRef":541,"sort":15,"status":16,"icon":384,"description":542,"httpMethod":64,"hasToc":27},"01a06b6a-0142-78fd-a603-078833fb6447","grp-api-bytedance-seedance-2-0-fast-image-to-video","video\u002Fbytedance\u002Fseedance-2-0-fast\u002Fimage-to-video","Seedance 2.0 Fast Image-to-Video","01a06b6b-45f9-757b-b5db-afd04194722b","Seedance 2.0 Fast Image-to-Video offers the highest image quality. It generates 4-15 second videos with text prompts and supports various aspect ratios, audio generation, and enhanced web search capabilities.",{"id":544,"groupId":545,"locale":4,"slug":546,"title":505,"pageType":17,"contentSource":24,"contentRef":547,"sort":15,"status":16,"icon":384,"description":548,"httpMethod":64,"hasToc":27},"01a06b6a-08cc-7ebc-a43c-d709711c0340","grp-api-bytedance-seedance-2-0","video\u002Fbytedance\u002Fseedance-2-0","01a06b6b-77a7-72ad-a282-d877c77c31fe","Seedance 2.0 offers the highest image quality. It generates 4-15 second videos with text prompts and supports various aspect ratios, audio generation, and enhanced web search capabilities.",{"id":550,"groupId":551,"locale":4,"slug":10,"title":552,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":384,"description":553,"hasToc":55,"children":554},"01a06b69-aa11-7cae-846a-61921f31a74f","grp-section-seedance2.0mini","Seedance 2.0 Mini","Seedance 2.0 Mini brings ByteDance’s advanced multimodal video generation architecture to velocity-driven and cost-sensitive production workflows. Designed as a lightweight, high-throughput model, Mini delivers the core capabilities of Seedance 2.0 at accelerated generation speeds and significantly reduced cost per video. Utilizing the exact same API endpoints and parameter schemas as the standard edition, developers can swap models with zero integration overhead.",[555,561,568],{"id":556,"groupId":557,"locale":4,"slug":558,"title":552,"pageType":17,"contentSource":24,"contentRef":559,"sort":15,"status":16,"icon":384,"description":560,"httpMethod":64,"hasToc":27},"01a06b69-fc3e-71b4-b9bb-2e6dcab72289","grp-api-bytedance-seedance-2-0-mini","video\u002Fbytedance\u002Fseedance-2-0-mini","01a06b6b-244d-7849-9f45-008b984af154","Seedance 2.0-mini offers rapid video generation. It generates 4-15 second videos with text prompts, supports various aspect ratios, audio generation, and enhanced web search capabilities.Reference image width must be between 300 px and 6,000 px.",{"id":562,"groupId":563,"locale":4,"slug":564,"title":565,"pageType":17,"contentSource":24,"contentRef":566,"sort":15,"status":16,"icon":384,"description":567,"httpMethod":64,"hasToc":27},"01a06b69-fec0-7096-9ea7-28c3c52e1335","grp-api-bytedance-seedance-2-0-mini-text-to-video","video\u002Fbytedance\u002Fseedance-2-0-mini\u002Ftext-to-video","Seedance 2.0 Mini Text-to-Video","01a06b6b-355e-7de5-83d0-638b1ba6b75c","Seedance 2.0 Mini Text-to-Video offers rapid video generation. It generates 4-15 second videos with text prompts, supports various aspect ratios, audio generation, and enhanced web search capabilities.Reference image width must be between 300 px and 6,000 px.",{"id":569,"groupId":570,"locale":4,"slug":571,"title":572,"pageType":17,"contentSource":24,"contentRef":573,"sort":15,"status":16,"icon":384,"description":574,"httpMethod":64,"hasToc":27},"01a06b6a-0218-724e-a8ee-e3311471da7f","grp-api-bytedance-seedance-2-0-mini-image-to-video","video\u002Fbytedance\u002Fseedance-2-0-mini\u002Fimage-to-video","Seedance 2.0 Mini Image-to-Video","01a06b6b-4b9f-7e96-a7fd-be89eb65d6d2","Seedance 2.0 Mini Image-to-Video offers rapid video generation. It generates 4-15 second videos with text prompts, supports various aspect ratios, audio generation, and enhanced web search capabilities.Reference image width must be between 300 px and 6,000 px.",{"id":576,"groupId":577,"locale":4,"slug":10,"title":578,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":579,"description":580,"hasToc":55,"children":581},"01a06b69-bd08-733a-862b-5a47c412ad1f","grp-section-happyhorse","Happy Horse","happyhorse","Gain production-grade API access to Alibaba’s upgraded AI video foundation model, Happy Horse 1.1. Building upon the benchmark-topping performance of version 1.0, Happy Horse 1.1 introduces smoother motion dynamics, improved temporal coherence, and hyper-realistic skin textures for close-up shots. It supporting up to 9 reference images for rigid character and style preservation. Rendering 3-to-15-second clips in up to 1080p resolution.",[582,589,596,603,610,617],{"id":583,"groupId":584,"locale":4,"slug":585,"title":586,"pageType":17,"contentSource":24,"contentRef":587,"sort":15,"status":16,"icon":579,"description":588,"httpMethod":64,"hasToc":27},"01a06b69-dae0-7f35-b59d-05b6ea9d7341","grp-api-ali-happyhorse-1-1-reference-to-video","video\u002Fali\u002Fhappyhorse-1-1\u002Freference-to-video","HappyHorse 1.1 Reference-to-Video","01a06b6a-637a-7195-bbe3-5c6ce9cc9b0b","HappyHorse 1.1 Reference-to-Video is Alibaba's next-generation AI reference-guided video generation model. Supporting 1 to 9 reference images (covering character identity, outfits, product silhouette, and scene styles), it achieves high-precision multi-image fusion and subject locking to prevent visual drift across shots. It generates native 720P\u002F1080P HD videos up to 15 seconds long per run.",{"id":590,"groupId":591,"locale":4,"slug":592,"title":593,"pageType":17,"contentSource":24,"contentRef":594,"sort":15,"status":16,"icon":579,"description":595,"httpMethod":64,"hasToc":27},"01a06b69-dbb6-7dac-8120-1184afb77a41","grp-api-ali-happyhorse-1-1-reference-to-video-spicy","video\u002Fali\u002Fhappyhorse-1-1\u002Freference-to-video-spicy","HappyHorse 1.1 Reference-to-Video Spicy","01a06b6a-689a-79a0-af56-96ec1a5bfaff","HappyHorse 1.1 Reference-to-Video Spicy is Alibaba's high-expressiveness, high-dynamism variant in the HappyHorse model series. Building upon multi-image subject locking and native audio-visual synchronization, the Spicy mode is a uncensored and fine-tuned for high-intensity action, aggressive camera tracking, strong visual impact, and dramatic visual effects. It leverages reference images to drive bold, highly expressive motion sequences and complex camera maneuvers while maintaining multilingual lip-sync and ambient audio—ideal for action-packed short dramas, viral social video, game CG FX, and high-impact commercial ads.",{"id":597,"groupId":598,"locale":4,"slug":599,"title":600,"pageType":17,"contentSource":24,"contentRef":601,"sort":15,"status":16,"icon":579,"description":602,"httpMethod":64,"hasToc":27},"01a06b69-dc8b-7401-a45e-d14ccc1fbbd0","grp-api-ali-happyhorse-1-1-image-to-video-spicy","video\u002Fali\u002Fhappyhorse-1-1\u002Fimage-to-video-spicy","HappyHorse 1.1 Image-to-Video Spicy","01a06b6a-6bf5-7f1a-9243-8386e2ca016d","HappyHorse 1.1 Image to Video Spicy transforms a single starting image into a short cinematic video, combining rock-solid temporal consistency with smooth, highly expressive character movement.",{"id":604,"groupId":605,"locale":4,"slug":606,"title":607,"pageType":17,"contentSource":24,"contentRef":608,"sort":15,"status":16,"icon":579,"description":609,"httpMethod":64,"hasToc":27},"01a06b69-dd5e-7e1c-9a8a-cccaf5ec9245","grp-api-ali-happyhorse-1-1-text-to-video-spicy","video\u002Fali\u002Fhappyhorse-1-1\u002Ftext-to-video-spicy","HappyHorse 1.1 Text-to-Video Spicy","01a06b6a-6f65-77b8-9d5e-baa3f9cc8133","HappyHorse 1.1 Text to Video Spicy turns simple text prompts into short cinematic clips, blending impressive temporal stability with expressive, nuanced character movement.",{"id":611,"groupId":612,"locale":4,"slug":613,"title":614,"pageType":17,"contentSource":24,"contentRef":615,"sort":15,"status":16,"icon":579,"description":616,"httpMethod":64,"hasToc":27},"01a06b69-f8e8-7d27-9578-8769e39c9133","grp-api-ali-happyhorse-1-1-text-to-video","video\u002Fali\u002Fhappyhorse-1-1\u002Ftext-to-video","HappyHorse 1.1 Text-to-Video","01a06b6b-0cdf-79a9-89bc-431bafe2985f","HappyHorse 1.1 Text-to-Video is Alibaba's next-generation AI text-to-video model. Built on an integrated audio-video joint generation architecture, it generates native 720P\u002F1080P HD videos directly from text, supporting up to 15 seconds of rendering. It natively supports multilingual lip-sync, ambient sound effects, and audio-visual synchronization without extra dubbing, while delivering exceptional motion smoothness, subject consistency, and camera control—ideal for short dramas, commercial ads, and social media video production.",{"id":618,"groupId":619,"locale":4,"slug":620,"title":621,"pageType":17,"contentSource":24,"contentRef":622,"sort":15,"status":16,"icon":579,"description":623,"httpMethod":64,"hasToc":27},"01a06b69-f9bb-7e3b-bf96-a9e64af4ae25","grp-api-ali-happyhorse-1-1-image-to-video","video\u002Fali\u002Fhappyhorse-1-1\u002Fimage-to-video","HappyHorse 1.1 Image-to-Video","01a06b6b-1286-7ce2-9783-17741e2b6eaf","HappyHorse 1.1 Image-to-Video is Alibaba's next-generation AI image-to-video model. Supporting first-frame driving, first-to-last frame transitions, and short video extension, it generates native 720P\u002F1080P HD videos up to 15 seconds per run. Built on an integrated audio-video joint generation architecture, it natively supports multilingual lip-sync, ambient sound effects, and audio-visual synchronization, while delivering exceptional motion smoothness, subject consistency, and camera control—ideal for e-commerce, short drama VFX, and social media production.",{"id":625,"groupId":626,"locale":4,"slug":10,"title":627,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":155,"description":628,"hasToc":55,"children":629},"01a06b69-b7d2-74b8-866e-7c56ee8c1cb2","grp-section-gemini-omni-flash","Gemini Omni Flash","Gain production-grade API access to Google’s flagship AI video generation and multi-turn editing model, Gemini Omni Flash. Reimagining the video generation stack, Omni Flash combines real-world physics simulation with rich spatial intelligence. It seamlessly transforms text prompts, multi-reference photos, audio, or existing video files into high-resolution videos complete with natively generated, synchronized audio.",[630,637],{"id":631,"groupId":632,"locale":4,"slug":633,"title":634,"pageType":17,"contentSource":24,"contentRef":635,"sort":15,"status":16,"icon":155,"description":636,"httpMethod":64,"hasToc":27},"01a079c7-9684-721d-91a1-5db801156448","grp-api-atlas-gemini-omni-flash-image-to-video","video\u002Fatlas\u002Fgemini-omni-flash\u002Fimage-to-video","Gemini Omni Flash Image-to-Video","01a06c19-23b7-74fe-bb38-2b3386df60c6","Gemini Omni Flash Image-to-Video is a next-generation multimodal video generation model developed by Google DeepMind. Built on a native Omni architecture, it accurately parses text prompts and input image semantics to produce cinematic 24 FPS dynamic videos. Supporting 16:9 and 9:16 aspect ratios, it generates 3–10 second fluid clips per run (featuring native support for up to 4K super-sampled upscaling, with 720P currently available on select platform endpoints). With exceptional subject consistency, physical simulation, and camera control, it excels in short-form drama, commercial advertising, film VFX, and social media animation.",{"id":638,"groupId":639,"locale":4,"slug":640,"title":641,"pageType":17,"contentSource":24,"contentRef":642,"sort":15,"status":16,"icon":155,"description":643,"httpMethod":64,"hasToc":27},"01a079c7-a4b3-7839-827c-eb21397acbd6","grp-api-atlas-gemini-omni-flash-text-to-video","video\u002Fatlas\u002Fgemini-omni-flash\u002Ftext-to-video","Gemini Omni Flash Text-to-Video","01a06c19-448e-7ac8-b1aa-2f57a8609c87","Gemini Omni Flash Text-to-Video is Google DeepMind's flagship unified multimodal model under the Gemini Omni series. Integrating Gemini's advanced reasoning with Veo's video generation capability, it enables native \"any-to-any\" generation, producing high-quality 24 FPS videos with synchronized audio directly from text prompts. It supports 16:9 and 9:16 aspect ratios, 3–10s video duration, and up to 1080P\u002F4K supersampled output. Featuring multi-turn conversational video editing with context retention and built-in SynthID digital watermarking, it is ideal for short-form video creation, commercial ads, film VFX, and multimodal Agent workflows.",{"id":645,"groupId":646,"locale":4,"slug":10,"title":647,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":266,"description":648,"hasToc":55,"children":649},"01a06b69-adb5-7123-ae76-8ca67cdb3b34","grp-section-minimax-h3","MiniMax H3","Moving beyond traditional single-task pipelines, MiniMax H3 leverages the H3-Omni Transformer architecture to deliver unified understanding and generation across interleaved text, image, video, and audio contexts. The model generates smooth, continuous video clips from 4 to 15 seconds at up to 2K crisp resolution, natively accompanied by 32kHz stereo audio including lip-synced speech, ambient foley, and soundtrack effects.",[650,657,663],{"id":651,"groupId":652,"locale":4,"slug":653,"title":654,"pageType":17,"contentSource":24,"contentRef":655,"sort":15,"status":16,"icon":266,"description":656,"httpMethod":64,"hasToc":27},"01a06b69-d3af-7399-b881-387756bc64c7","grp-api-minimax-h3-video-image-to-video","video\u002Fminimax\u002Fh3-video\u002Fimage-to-video","MiniMax H3 Image-to-Video","01a06b6a-3a80-779b-87c5-08855740f7e7","MiniMax H3 Image-to-Video is MiniMax's next-generation multimodal AI video model. Supporting first-frame driving and first-to-last frame transitions, it generates up to 2K cinematic HD videos directly, with durations ranging from 5 to 15 seconds. Built on a unified Omni architecture, it natively supports integrated audio-video generation (sound effects, ambient audio, and multilingual lip-sync) alongside exceptional camera control, physics simulation, and subject consistency—ideal for e-commerce, commercial ads, and short drama production.",{"id":658,"groupId":659,"locale":4,"slug":660,"title":647,"pageType":17,"contentSource":24,"contentRef":661,"sort":15,"status":16,"icon":266,"description":662,"httpMethod":64,"hasToc":27},"01a06b69-d4e3-778e-b47a-392366d012e3","grp-api-minimax-h3-video","video\u002Fminimax\u002Fh3-video","01a06b6a-3fe4-7ed4-98a9-abe9f3d8efc2","MiniMax H3 : generate a video that keeps the subject from a reference image, driven by a text prompt. Supports 2K, 5-15s.",{"id":664,"groupId":665,"locale":4,"slug":666,"title":667,"pageType":17,"contentSource":24,"contentRef":668,"sort":15,"status":16,"icon":266,"description":669,"httpMethod":64,"hasToc":27},"01a06b69-d84f-72d8-b4b0-14cadcded42c","grp-api-minimax-h3-video-text-to-video","video\u002Fminimax\u002Fh3-video\u002Ftext-to-video","MiniMax H3 Text-to-Video","01a06b6a-557b-739f-801f-cff089309b48","MiniMax H3 Text-to-Video is MiniMax's next-generation AI video generation model. Powered by a unified Omni architecture, it accurately parses complex prompt text to directly generate up to 2K cinematic-grade videos up to 15 seconds long. It natively supports integrated audio-video generation (ambient audio, sound effects, and multilingual lip-sync) alongside exceptional motion smoothness, physical simulation, and camera control—ideal for commercial advertising, short dramas, and social media video creation.",{"id":671,"groupId":672,"locale":4,"slug":10,"title":673,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":674,"description":675,"hasToc":55,"children":676},"01a06b69-a5bd-737e-89a7-b750f4c6ee6e","grp-section-wan3.0","Wan 3.0","wan","Wan 3.0 is Alibaba's next-generation multimodal AI video generation model. Supporting both text-to-video and image-to-video workflows, it generates up to 30-second videos in up to 1080P\u002F4K cinematic resolution.",[677,684,691,698,705,712,719,726,733,740,747,754],{"id":678,"groupId":679,"locale":4,"slug":680,"title":681,"pageType":17,"contentSource":24,"contentRef":682,"sort":15,"status":16,"icon":674,"description":683,"httpMethod":64,"hasToc":27},"01a06b69-e03e-7705-8416-e1301fda92da","grp-api-aliyun-wan3-0-image-to-video","video\u002Faliyun\u002Fwan3-0\u002Fimage-to-video","Wan 3.0 Image-to-Video","01a06b6a-7f39-7816-ad35-617fbd596a2b","Wan 3.0 Image-to-Video is Alibaba's next-generation AI image-to-video model under the Tongyi Wanxiang series. Supporting single first-frame driving and smooth first-to-last frame transitions, it directly generates up to 30-second videos in up to 1080P\u002F4K cinematic resolution.",{"id":685,"groupId":686,"locale":4,"slug":687,"title":688,"pageType":17,"contentSource":24,"contentRef":689,"sort":15,"status":16,"icon":674,"description":690,"httpMethod":64,"hasToc":27},"01a06b69-e2c9-73b5-bf8a-99426f8f9440","grp-api-aliyun-wan3-0-image-to-video-global","video\u002Faliyun\u002Fwan3-0\u002Fimage-to-video-global","Wan 3.0 Image-to-Video Spicy","01a06b6a-8da2-7526-a544-c2e48f4d1a4c","Wan 3.0 Image-to-Video Spicy is Alibaba's high-dynamic image-to-video model under the Tongyi Wanxiang framework. Specially engineered for large-scale motion and high visual intensity, it transforms a single static image into up to 30-second videos in cinematic 1080P resolution. While maintaining strict character identity and background consistency from the input image, the \"Spicy\" edition delivers a dramatic boost in action amplitude, complex physical collision simulation, and expressive camera movement, alongside native audio-visual generation capabilities. It is ideally suited for high-energy commercial advertising, game CG animation, short drama production, and advanced visual effects workflows.",{"id":692,"groupId":693,"locale":4,"slug":694,"title":695,"pageType":17,"contentSource":24,"contentRef":696,"sort":15,"status":16,"icon":674,"description":697,"httpMethod":64,"hasToc":27},"01a06b69-e3a4-79e0-bfe9-3b25f72700ff","grp-api-aliyun-wan3-0-prime-image-to-video-global","video\u002Faliyun\u002Fwan3-0-prime\u002Fimage-to-video-global","Wan 3.0 Prime Image-to-Video Spicy","01a06b6a-90fc-75a0-b695-93a527a61f43","Wan 3.0 Prime Image-to-Video Spicy is Alibaba's high-speed, high-expressiveness AI image-to-video model variant under the Tongyi Wanxiang family. Combining Prime's ultra-fast generation inference with Spicy's high-dynamism visual tuning, it animates source images (with optional end-frame guidance) into up to 30-second 1080P HD videos in a single pass.",{"id":699,"groupId":700,"locale":4,"slug":701,"title":702,"pageType":17,"contentSource":24,"contentRef":703,"sort":15,"status":16,"icon":674,"description":704,"httpMethod":64,"hasToc":27},"01a06b69-e4e2-7d68-bc5c-29a4e1f21811","grp-api-aliyun-wan3-0-text-to-video-global","video\u002Faliyun\u002Fwan3-0\u002Ftext-to-video-global","Wan 3.0 Text-to-Video Spicy","01a06b6a-9466-7dd1-8d22-2a9303380e61","Wan 3.0 Text-to-Video Spicy is Alibaba's high-expressiveness AI text-to-video model variant under the Tongyi Wanxiang series. Building upon the core Wan 3.0 architecture, the Spicy edition is fine-tuned for high-intensity physical motion, dramatic camera maneuvering, high-contrast lighting, and powerful visual impact, directly generating up to 30-second videos in up to 1080P cinematic resolution.",{"id":706,"groupId":707,"locale":4,"slug":708,"title":709,"pageType":17,"contentSource":24,"contentRef":710,"sort":15,"status":16,"icon":674,"description":711,"httpMethod":64,"hasToc":27},"01a06b69-e5b6-7adb-bf2d-a31e1e21bcd2","grp-api-aliyun-wan3-0-prime-text-to-video-global","video\u002Faliyun\u002Fwan3-0-prime\u002Ftext-to-video-global","Wan 3.0 Prime Text-to-Video Spicy","01a06b6a-9815-71b0-8ca6-66d662747378","Wan 3.0 Prime Text-to-Video Spicy is Alibaba's high-speed, high-expressiveness AI text-to-video model variant under the Tongyi Wanxiang family. Combining Prime's ultra-fast generation inference with Spicy's high-dynamism visual tuning, it deeply parses complex text prompts to directly render up to 30-second 1080P HD videos with significantly reduced wait times. Fine-tuned for bold physical movement, high-contrast lighting, dramatic camera maneuvers, and intense visual impact, it natively supports integrated audio-video generation (ambient audio, sound effects, and multilingual lip-sync)—delivering rapid turnarounds and extreme visual tension for high-energy social media content, action sequences, and commercial ads.",{"id":713,"groupId":714,"locale":4,"slug":715,"title":716,"pageType":17,"contentSource":24,"contentRef":717,"sort":15,"status":16,"icon":674,"description":718,"httpMethod":64,"hasToc":27},"01a06b69-e68c-7718-89d0-76bd2a98baa4","grp-api-aliyun-wan3-0-text-to-video","video\u002Faliyun\u002Fwan3-0\u002Ftext-to-video","Wan 3.0 Text-to-Video","01a06b6a-9d53-7a2d-8a3a-92d187f7055e","Wan 3.0 Text-to-Video is Alibaba's next-generation AI text-to-video model under the Tongyi Wanxiang series. It deeply parses complex prompt text to directly generate up to 30-second videos in up to 1080P\u002F4K cinematic resolution. Featuring native audio-visual synchronization (ambient audio, sound effects, and multilingual lip-sync), it delivers exceptional motion smoothness, physical simulation, and precise camera control—ideal for commercial advertising, short dramas, film VFX, and social media content creation.",{"id":720,"groupId":721,"locale":4,"slug":722,"title":723,"pageType":17,"contentSource":24,"contentRef":724,"sort":15,"status":16,"icon":674,"description":725,"httpMethod":64,"hasToc":27},"01a08015-ac03-7618-9ef5-92419ecccfd8","grp-api-aliyun-wan3-0-prime-text-to-video","video\u002Faliyun\u002Fwan3-0-prime\u002Ftext-to-video","Wan 3.0 Prime Text-to-Video","01a08017-aa7c-7621-82e2-74f2d8b80caa","Wan 3.0 Prime Text-to-Video is Alibaba's high-speed AI text-to-video model under the Tongyi Wanxiang family. Combining the Prime architecture's rapid inference with the core Wan 3.0 multimodal foundation, it deeply parses complex text prompts to directly render up to 30-second 1080P HD videos with significantly reduced generation latency. Featuring native audio-visual synchronization (ambient audio, sound effects, and multilingual lip-sync), it delivers exceptional physical motion simulation, seamless temporal coherence, and precise camera control—providing rapid turnarounds and high-quality visual output for commercial advertising, short dramas, film VFX, and high-frequency social media content creation.",{"id":727,"groupId":728,"locale":4,"slug":729,"title":730,"pageType":17,"contentSource":24,"contentRef":731,"sort":15,"status":16,"icon":674,"description":732,"httpMethod":64,"hasToc":27},"01a08015-acdc-7cc9-b456-309bc8c83c25","grp-api-aliyun-wan3-0-prime-image-to-video","video\u002Faliyun\u002Fwan3-0-prime\u002Fimage-to-video","Wan 3.0 Prime Image-to-Video","01a08017-c742-7c58-abba-b38c3bed55ce","Wan 3.0 Prime Image-to-Video is Alibaba's high-speed AI image-to-video model under the Tongyi Wanxiang family. Leveraging the Prime architecture's rapid inference alongside the Wan 3.0 multimodal foundation, it animates source images (with optional end-frame guidance) into up to 30-second 1080P HD videos with significantly reduced rendering latency. Featuring native audio-visual synchronization (ambient audio, sound effects, and multilingual lip-sync), it delivers accurate physical motion simulation, precise camera control, and strong subject consistency—ideal for fast-turnaround e-commerce animation, film VFX, short dramas, and commercial advertising.",{"id":734,"groupId":735,"locale":4,"slug":736,"title":737,"pageType":17,"contentSource":24,"contentRef":738,"sort":15,"status":16,"icon":674,"description":739,"httpMethod":64,"hasToc":27},"01a0840c-97aa-7d05-9954-90e7e1fba081","grp-api-aliyun-wan3-0-prime-reference-to-video-global","video\u002Faliyun\u002Fwan3-0-prime\u002Freference-to-video-global","Wan 3.0 Prime Reference-to-Video Spicy","01a08072-2066-79a6-9228-ad2c393eef42","Wan 3.0 Prime Reference-to-Video Spicy is Alibaba's flagship uncensored high-motion video model engineered for native 4K visual synthesis. Delivering native 4K output at 60 fps, it processes up to 8 reference images or 2 reference videos simultaneously to preserve character identity and material fidelity across complex sequences. With unrestricted generation capabilities, it synthesizes explosive physical movements, dramatic camera trajectories, and physical interactions—powering high-energy cinematic pre-visualization, high-impact commercial ads, AAA gaming assets, and action pipelines.",{"id":741,"groupId":742,"locale":4,"slug":743,"title":744,"pageType":17,"contentSource":24,"contentRef":745,"sort":15,"status":16,"icon":674,"description":746,"httpMethod":64,"hasToc":27},"01a0840c-9885-7523-98ae-88116a4c8770","grp-api-aliyun-wan3-0-reference-to-video-global","video\u002Faliyun\u002Fwan3-0\u002Freference-to-video-global","Wan 3.0 Reference-to-Video Spicy","01a08072-23db-7d0a-93f4-fb743dc1076d","Wan 3.0 Reference-to-Video Spicy is Alibaba's high-motion video synthesis model engineered for high-energy visual dynamics and stylized action generation. Built on an expanded spatio-temporal attention mechanism alongside robust reference feature anchoring, it synthesizes large-scale physical movements, aggressive camera trajectories, and dramatic temporal transitions while preserving strict subject identity and garment texture fidelity. Wan 3.0 Reference-to-Video Spicy powers dynamic action video production, cinematic FX pre-visualization, interactive gaming visual assets, and high-impact commercial advertising pipelines.",{"id":748,"groupId":749,"locale":4,"slug":750,"title":751,"pageType":17,"contentSource":24,"contentRef":752,"sort":15,"status":16,"icon":674,"description":753,"httpMethod":64,"hasToc":27},"01a0840c-9960-7207-a998-16cb9f26ad3b","grp-api-aliyun-wan3-0-prime-reference-to-video","video\u002Faliyun\u002Fwan3-0-prime\u002Freference-to-video","Wan 3.0 Prime Reference-to-Video","01a08072-2972-702f-b714-9c45c1847434","Wan 3.0 Prime Reference-to-Video is Alibaba's flagship controllable video generation model built for precise reference-conditioned visual synthesis. Powered by an upgraded spatio-temporal decoupling architecture and multi-reference feature fusion, it preserves character identity, garment textures, ambient lighting, and complex camera trajectories across extended sequences while generating native 4K high-frame-rate video. Operating with strict temporal consistency and fluid motion dynamics, Wan 3.0 Prime drives e-commerce video production, cinematic pre-visualization, digital human animation, and commercial advertising pipelines.",{"id":755,"groupId":756,"locale":4,"slug":757,"title":758,"pageType":17,"contentSource":24,"contentRef":759,"sort":15,"status":16,"icon":674,"description":760,"httpMethod":64,"hasToc":27},"01a0840c-9a42-7d6f-b67f-907ecfb51fbf","grp-api-aliyun-wan3-0-reference-to-video","video\u002Faliyun\u002Fwan3-0\u002Freference-to-video","Wan 3.0 Reference-to-Video","01a08072-2ec5-7807-8996-413c807956e8","Wan 3.0 Reference-to-Video is Alibaba's flagship high-efficiency video generation model engineered for precise reference-driven visual synthesis. Built on an upgraded multi-modal reference feature alignment and spatio-temporal attention architecture, it preserves strict subject identity, garment textures, and ambient lighting across dynamic video sequences. Combining high-throughput inference speeds with fluid camera movement, Wan 3.0 Reference-to-Video efficiently powers e-commerce product videos, digital human animation, social media marketing assets, and commercial advertising pipelines.",{"id":762,"groupId":763,"locale":4,"slug":10,"title":764,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":674,"description":765,"hasToc":55,"children":766},"01a06b69-b3bd-7c51-9ac0-9f3b2ab58a0a","grp-section-wan2.7","Wan 2.7","Gain production-grade API access to Alibaba Tongyi Lab’s flagship video generation suite, Wan 2.7. Designed to empower professional production pipelines with unprecedented directorial control, Wan 2.7 leverages a 27B MoE architecture that bridges generation and post-production. Beyond high-fidelity text-to-video and image-to-video with synchronized audio, Wan 2.7 introduces pioneering first-and-last-frame interpolation and instruction-based video-to-video editing.",[767,774,781,788],{"id":768,"groupId":769,"locale":4,"slug":770,"title":771,"pageType":17,"contentSource":24,"contentRef":772,"sort":15,"status":16,"icon":231,"description":773,"httpMethod":64,"hasToc":27},"01a06b69-d1e3-7047-bba3-632212aab205","grp-api-aliyun-wan2-7-text-to-video","video\u002Faliyun\u002Fwan2-7\u002Ftext-to-video","Wan 2.7 Text-to-Video","01a06b6a-2f25-7e33-84be-9f98b05e080c","Tongyi Wanxiang Wan 2.7 Text-to-Video is Alibaba Cloud's next-generation text-to-video model. Supporting long Chinese and English prompts with smart expansion, it generates native 720P\u002F1080P HD videos directly from text, up to 15 seconds per run. Featuring native audio-visual coordination and sound effect sync, it adaptively supports multiple aspect ratios like 16:9 and 9:16 with strong motion continuity, realistic physics simulation, and lighting rendering—ideal for commercial ads, short videos, and anime creation.",{"id":775,"groupId":776,"locale":4,"slug":777,"title":778,"pageType":17,"contentSource":24,"contentRef":779,"sort":15,"status":16,"icon":231,"description":780,"httpMethod":64,"hasToc":27},"01a06b69-d925-79bd-8b7d-3d0a39c4559b","grp-api-aliyun-wan2-7-text-to-video-sp","video\u002Faliyun\u002Fwan2-7\u002Ftext-to-video-sp","Wan 2.7 Text-to-Video Spicy","01a06b6a-58df-78b8-bd94-38ee40a1427d","Wan 2.7 Text-to-Video Spicy turns simple text prompts into short cinematic clips, blending impressive temporal stability with expressive, nuanced character movement.",{"id":782,"groupId":783,"locale":4,"slug":784,"title":785,"pageType":17,"contentSource":24,"contentRef":786,"sort":15,"status":16,"icon":231,"description":787,"httpMethod":64,"hasToc":27},"01a06b69-da00-7238-b6cc-9e589c409f73","grp-api-aliyun-wan2-7-image-to-video-sp","video\u002Faliyun\u002Fwan2-7\u002Fimage-to-video-sp","Wan 2.7 Image-to-Video Spicy","01a06b6a-5e68-78df-8d64-9e1b3e370c09","Wan 2.7 Image-to-Video Spicy turns a first-frame image into short cinematic motion with stable temporal detail and expressive character movement.",{"id":789,"groupId":790,"locale":4,"slug":791,"title":792,"pageType":17,"contentSource":24,"contentRef":793,"sort":15,"status":16,"icon":231,"description":794,"httpMethod":64,"hasToc":27},"01a06b69-f4b1-7fbd-8e95-d05306c8139e","grp-api-aliyun-wan2-7-image-to-video","video\u002Faliyun\u002Fwan2-7\u002Fimage-to-video","Wan 2.7 Image-to-Video","01a06b6a-f256-7a42-8c58-38537f6e6176","Tongyi Wanxiang Wan 2.7 Image-to-Video is Alibaba Cloud's next-generation image-to-video model. It supports first-frame generation, first-to-last frame transitions, and short video extension, generating 720P\u002F1080P HD videos up to 15 seconds per run. Featuring powerful camera control and realistic physics simulation, it natively supports audio-driven lip-sync and action alignment while adaptively supporting mainstream aspect ratios like 16:9 and 9:16—ideal for e-commerce, VFX, and film post-production.",{"id":796,"groupId":797,"locale":4,"slug":10,"title":798,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":799,"description":800,"hasToc":55,"children":801},"01a06b69-b4e2-75ff-946d-6141a2643e05","grp-section-kling3.0","Kling V3.0","kling","Gain production-grade API access to Kuaishou's flagship audiovisual generation suite, Kling 3.0. Built on a unified, native multimodal (All-in-One) training architecture, Kling 3.0 seamlessly integrates text-to-video, image-to-video, reference conditioning, and in-video editing into a streamlined workflow. Breaking previous video length barriers, it generates continuous 3 to 15-second cinematic sequences with native intelligent multi-shot storyboarding.",[802,809,816,823],{"id":803,"groupId":804,"locale":4,"slug":805,"title":806,"pageType":17,"contentSource":24,"contentRef":807,"sort":15,"status":16,"icon":799,"description":808,"httpMethod":64,"hasToc":27},"01a06b69-d2b7-78e2-8f2e-d4dede47c838","grp-api-kwaivgi-kling-v3-omni","video\u002Fkwaivgi\u002Fkling-v3-omni","Kling 3.0 Omni","01a06b6a-34cc-7fda-8c9a-fb20426e47b6","Kling 3.0 Omni delivers high-quality text-to-video generation with smooth motion, cinematic visuals, accurate prompt adherence, and native audio for ready-to-share clips. Ready-to-use REST inference API, best performance, no cold starts, affordable pricing.",{"id":810,"groupId":811,"locale":4,"slug":812,"title":813,"pageType":17,"contentSource":24,"contentRef":814,"sort":15,"status":16,"icon":799,"description":815,"httpMethod":64,"hasToc":27},"01a06b69-d5ba-7790-a1d7-316fe64059a4","grp-api-kwaivgi-kling-video-o1","video\u002Fkwaivgi\u002Fkling-video-o1","Kling Video O1","01a06b6a-458f-7395-abdb-5fbabc319ed9","Kling Video O1 (Omni One) is Kuaishou's industry-first unified multimodal video model that merges video generation and editing into a single engine. Integrating text-to-video, image-to-video, element referencing, localized inpainting, and video restyling, it supports referencing up to 7 subjects simultaneously to lock character and prop consistency. Creators can perform conversational video editing via natural language prompts—ideal for advertising, VFX, and end-to-end film production.",{"id":817,"groupId":818,"locale":4,"slug":819,"title":820,"pageType":17,"contentSource":24,"contentRef":821,"sort":15,"status":16,"icon":799,"description":822,"httpMethod":64,"hasToc":27},"01a06b69-f585-7153-be51-046f7796faa8","grp-api-kwaivgi-kling-v3-0-image-to-video","video\u002Fkwaivgi\u002Fkling-v3-0\u002Fimage-to-video","Kling v3.0 Image-to-Video","01a06b6a-f7cf-7b76-80fe-6d69009fd1a0","Kling v3.0 Image-to-Video is Kuaishou's next-generation multimodal AI video model. Built on the unified Omni architecture, it takes static images or subject references to generate up to 15-second cinematic videos in up to 4K resolution. It features native audio-visual synchronization, multilingual lip-sync, and enhanced subject consistency to prevent visual drift, along with intelligent multi-shot control—ideal for commercial ads, film VFX, and narrative short videos.",{"id":824,"groupId":825,"locale":4,"slug":826,"title":827,"pageType":17,"contentSource":24,"contentRef":828,"sort":15,"status":16,"icon":799,"description":829,"httpMethod":64,"hasToc":27},"01a06b69-fb62-7b62-b22d-0bad3bbcb59b","grp-api-kwaivgi-kling-v3-0-text-to-video","video\u002Fkwaivgi\u002Fkling-v3-0\u002Ftext-to-video","Kling v3.0 Text-to-Video","01a06b6b-1ec9-7cf6-9722-d0b0007c2d93","Kling v3.0 Text-to-Video is Kuaishou's next-generation AI video model. Powered by the native Omni architecture, it accurately parses complex prompt text to generate up to 4K cinematic-grade videos up to 15 seconds long. It natively supports integrated audio-video generation (ambient audio, music, and multilingual lip-sync) alongside exceptional visual realism, multi-shot coherence, and complex physical simulation—ideal for commercial advertising, film VFX, and content creation.",{"id":831,"groupId":832,"locale":4,"slug":10,"title":833,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"topTab":834,"children":835},"01a06b0c-7a2c-7126-bf27-c7e4296f6457","grp-section-chat-apps","Chat Apps","integration",[836,844,852,860,868,876],{"id":837,"groupId":838,"locale":4,"slug":839,"title":840,"pageType":32,"contentSource":24,"contentRef":841,"sort":15,"status":16,"icon":842,"description":843,"hasToc":27},"01a06b0c-7fdb-72fe-bebc-ac5ff942b35f","grp-tools-workbuddy","tools\u002Fworkbuddy","Workbuddy","01a06b0d-c1d1-7ffe-9b43-9712b1c672bb","codebuddy","Connect iCreat in Workbuddy",{"id":845,"groupId":846,"locale":4,"slug":847,"title":848,"pageType":32,"contentSource":24,"contentRef":849,"sort":15,"status":16,"icon":850,"description":851,"hasToc":27},"01a06b0c-80af-7ed4-83e3-cdf20f731954","grp-tools-codex","tools\u002Fcodex","Codex","01a06b0d-ccfa-7770-9f81-af321617b765","codex","Configure Codex to use iCreat",{"id":853,"groupId":854,"locale":4,"slug":855,"title":856,"pageType":32,"contentSource":24,"contentRef":857,"sort":15,"status":16,"icon":858,"description":859,"hasToc":27},"01a06b0c-8183-722e-b530-6ae8401757e2","grp-tools-claude-code","tools\u002Fclaude-code","Claude Code","01a06b0d-d764-7a47-8846-0af796ee9e70","claudecode","Configure Claude Code desktop app with iCreat",{"id":861,"groupId":862,"locale":4,"slug":863,"title":864,"pageType":32,"contentSource":24,"contentRef":865,"sort":15,"status":16,"icon":866,"description":867,"hasToc":27},"01a06b0c-82a7-7312-bb44-1a53216c8784","grp-tools-chatbox","tools\u002Fchatbox","ChatBox","01a06b0d-e170-7297-98f9-f81451631a3d","chatbox","Connect iCreat in ChatBox via OpenAI-compatible API",{"id":869,"groupId":870,"locale":4,"slug":871,"title":872,"pageType":32,"contentSource":24,"contentRef":873,"sort":15,"status":16,"icon":874,"description":875,"hasToc":27},"01a06b0c-837f-708e-9630-12ad12c05cc5","grp-tools-cherry-studio","tools\u002Fcherry-studio","Cherry Studio","01a06b0d-eba5-7e7f-a1c5-e1a46b613625","cherry-studio","Connect iCreat in Cherry Studio via OpenAI-compatible API",{"id":877,"groupId":878,"locale":4,"slug":879,"title":880,"pageType":32,"contentSource":24,"contentRef":881,"sort":15,"status":16,"icon":882,"description":883,"hasToc":27},"01a06b0c-8451-7c8b-935f-0eb7a109d461","grp-tools-anythingllm","tools\u002Fanythingllm","AnythingLLM","01a06b0d-f593-74d3-85fb-78dea17a7a41","anythingllm","Connect iCreat in AnythingLLM via OpenAI-compatible API",{"id":885,"groupId":886,"locale":4,"slug":10,"title":887,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"topTab":834,"children":888},"01a06b0c-7b01-74c7-8611-ffe1c08e2cac","grp-section-dev-tools","Dev Tools",[889,896,904,912],{"id":890,"groupId":891,"locale":4,"slug":892,"title":893,"pageType":32,"contentSource":24,"contentRef":894,"sort":15,"status":16,"icon":858,"description":895,"hasToc":27},"01a06b0c-8524-7d51-94c9-d02f862eb2d9","grp-tools-claude-code-cli","tools\u002Fclaude-code-cli","Claude Code CLI","01a06b0e-0065-7fc5-8619-e5d16b1bf760","Connect iCreat via Claude Code CLI in terminal",{"id":897,"groupId":898,"locale":4,"slug":899,"title":900,"pageType":32,"contentSource":24,"contentRef":901,"sort":15,"status":16,"icon":902,"description":903,"hasToc":27},"01a06b0c-85fb-7a48-bc35-27a8b0cb79f1","grp-tools-cursor","tools\u002Fcursor","Cursor","01a06b0e-0b36-77a5-9c70-2e7433b05e5c","cursor","Connect iCreat API in Cursor editor",{"id":905,"groupId":906,"locale":4,"slug":907,"title":908,"pageType":32,"contentSource":24,"contentRef":909,"sort":15,"status":16,"icon":910,"description":911,"hasToc":27},"01a06b0c-86cf-788f-a18e-87a908200893","grp-tools-opencode","tools\u002Fopencode","OpenCode","01a06b0e-1649-776a-b701-9bc8061c0f71","opencode","Connect iCreat API in OpenCode",{"id":913,"groupId":914,"locale":4,"slug":915,"title":916,"pageType":32,"contentSource":24,"contentRef":917,"sort":15,"status":16,"icon":918,"description":919,"hasToc":27},"01a06b0c-87a3-7758-b380-6fd09cc1cc6c","grp-tools-cc-switch","tools\u002Fcc-switch","CC-Switch","01a06b0e-2170-76b1-b084-da669f65f2d8","cc-switch","Manage CLI configs with CC-Switch",{"id":921,"groupId":922,"locale":4,"slug":10,"title":923,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"topTab":924,"children":925},"01a06b0c-7c21-78b9-9d90-2a1e540fcd59","grp-section-help","Help","faq",[926],{"id":927,"groupId":928,"locale":4,"slug":924,"title":929,"pageType":32,"contentSource":24,"contentRef":930,"sort":15,"status":16,"icon":931,"description":932,"hasToc":27},"01a06b0c-8878-7b26-9746-ebe7d6a8f76b","grp-faq","FAQ","01a06b0e-2cd3-714b-be93-99c0f5af71fc","line-search","Frequently asked questions",{"locale":934,"updatedAt":935,"items":936},"zh","2026-09-20T05:24:50.689592266Z",[937,954,1132,1220,1405,1433,1453],{"id":938,"groupId":9,"locale":934,"slug":10,"title":939,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"topTab":17,"children":940},"01a06b0c-62e7-7e35-9aa1-8ab2a9f8cba3","快速开始",[941,946,950],{"id":942,"groupId":21,"locale":934,"slug":22,"title":943,"pageType":12,"contentSource":24,"contentRef":944,"sort":15,"status":16,"icon":12,"description":945,"hasToc":27},"01a06b0c-6920-77d3-a990-7e754f389267","概览","01a06b0d-9d7d-7636-9e01-cd823e53fb38","iCreat AI 是一站式生成式 AI 模型服务平台",{"id":947,"groupId":30,"locale":934,"slug":31,"title":939,"pageType":32,"contentSource":24,"contentRef":948,"sort":15,"status":16,"icon":34,"description":949,"hasToc":27},"01a06b0c-6a14-7311-8095-17657b339ab8","01a06b0d-a872-7880-ae04-e5dd29dc3dbd","从注册账户到完成首次 API 调用的完整接入指南",{"id":951,"groupId":38,"locale":934,"slug":39,"title":40,"pageType":32,"contentSource":24,"contentRef":952,"sort":15,"status":16,"icon":42,"description":953,"hasToc":27},"01a06b0c-6aec-7cd9-8c4f-2027cf6f70ef","01a06b0d-b2ac-735c-915d-3e2e5a31080e","创建和管理用于身份验证的 API Key",{"id":955,"groupId":46,"locale":934,"slug":10,"title":47,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"topTab":17,"children":956},"01a06b0c-63c4-7b5f-9f17-ccad331d2a26",[957,1001,1025,1041,1057,1069,1081,1089,1109,1124],{"id":958,"groupId":51,"locale":934,"slug":10,"title":52,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":53,"description":959,"hasToc":55,"children":960},"01a06b45-46df-72f5-a1f1-1c53aadf06e4","在全新一代架构中，Claude 家族构建了清晰的分级能力矩阵：包括能力上限极高的 Claude Fable 5（前沿深层推演与长路线 Agent）、专为复杂软件工程与企业级代码管线设计的 Claude Opus 5、兼具速度与强逻辑的 Claude Sonnet 5，以及高吞吐低延迟的 Claude Haiku 4.5。Claude 5 套件全面支持五阶动态思考调控（Low\u002FMedium\u002FHigh\u002FXHigh\u002FMax Thinking Effort），并具备 1,000,000 Tokens 超长无损上下文与 Prompt Caching（缓存）降本机制。",[961,966,971,976,981,986,991,996],{"id":962,"groupId":59,"locale":934,"slug":60,"title":963,"pageType":17,"contentSource":24,"contentRef":964,"sort":15,"status":16,"icon":53,"description":965,"httpMethod":64,"hasToc":27},"01a08015-0d1e-7e3b-a600-9773672eb16d","Claude Fable 5.1 经济版","01a08017-4779-77f4-905f-8b070087fc0c","Claude Fable 5.1 是 Anthropic 推出的新一代前沿级旗舰 AI 大模型，针对高难度代码工程、长流程 Agent 自主协作与复杂知识工作进行了深度升级。模型具备 Mythos 级别的强推理能力，强化了自主自我验证与故障根因排查（Root-cause Fixing），并将安全拦截假阳性率降低达 60%。5.1 版本大幅优化了 Prompt 缓存架构（Cache Read 成本降低 75%），使高 Agent 化任务的综合运行成本降低高达 45%。模型原生支持软件漏洞挖掘与隐形水印检测技术，广泛适用于全流程软件工程、高级金融建模分析、企业级网络安全审计与长程科研任务。",{"id":967,"groupId":67,"locale":934,"slug":68,"title":968,"pageType":17,"contentSource":24,"contentRef":969,"sort":15,"status":16,"icon":53,"description":970,"httpMethod":64,"hasToc":27},"01a06b45-a2e2-7b3e-9c84-c806d62a12ae","Claude Fable 5 经济版","01a06b6a-8707-7428-a312-5d0d48c85e74","Claude Fable 5 是 Anthropic 正式发布的 Mythos 级模型，专为自主知识工作、高级推理和长周期编程而设计。它支持文本、图像和文件输入，并生成文本输出。该模型拥有 100 万 tokens 的上下文窗口、最多 128,000 个输出 tokens，并内置始终启用的自适应思考、工具使用和结构化输出能力。它与 Claude Mythos 5 属于同一底层模型家族，但公开发布时采用了更强的安全防护措施。该模型适用于复杂软件工程、多步骤智能体工作流、企业研究、文档分析，以及需要持续推理和高可靠性的专业任务。",{"id":972,"groupId":74,"locale":934,"slug":75,"title":973,"pageType":17,"contentSource":24,"contentRef":974,"sort":15,"status":16,"icon":53,"description":975,"httpMethod":64,"hasToc":27},"01a06b45-c562-77c6-a49c-4813b1a37894","Claude Sonnet 5 经济版","01a06b6b-5428-7782-ae8f-bd53a55525df","Sonnet 5 是 Anthropic 能力最强的 Sonnet 级模型，在编程、智能体工作流和专业任务方面具备前沿水平的性能。它支持自适应思考，并提供 **low**、**medium**、**high**、**max** 和 **x-high** 等可选推理级别，同时拥有 100 万 tokens 的上下文窗口，支持文本、图像和文件输入。\n\nSonnet 5 采用更新后的分词器，并集成实时网络安全防护机制，用于阻止高风险的军民两用活动。",{"id":977,"groupId":81,"locale":934,"slug":82,"title":978,"pageType":17,"contentSource":24,"contentRef":979,"sort":15,"status":16,"icon":53,"description":980,"httpMethod":64,"hasToc":27},"01a08015-0c4e-7adb-bdd0-53ec504b8b43","Claude Opus 5 经济版","01a08017-423d-779b-8c40-76d8ebedda24","Claude Opus 5 是 Anthropic 推出的新一代超旗舰级大模型，代表了当前 AI 在深度逻辑推理、高难度科学合成与复杂架构设计领域的最高智能水准。模型原生支持 100 万 Token 上下文与极致的自适应思考机制，能在多步推演、高级编程与跨学科科研任务中展现出接近人类顶级专家的分析能力与极低的幻觉率。结合企级安全对齐机制，Claude Opus 5 广泛适用于高难度法律合规审查、复杂金融衍生品建模、前沿生物医药研发与零容错的核心系统架构审计。",{"id":982,"groupId":88,"locale":934,"slug":89,"title":983,"pageType":17,"contentSource":24,"contentRef":984,"sort":15,"status":16,"icon":53,"description":985,"httpMethod":64,"hasToc":27},"01a06b45-c63c-767f-b2b0-076d89769a4c","Claude Opus 4.8 经济版","01a06b6b-595f-76c9-a5fc-178a98d35c08","Claude Opus 4.8 是 Anthropic 正式发布的最强 Opus 模型，专为复杂推理、长周期智能体编程和高度自主的专业工作流而设计。它支持文本、图像和文件输入，并生成文本输出。该模型拥有 100 万 tokens 的上下文窗口、最多 128,000 个输出 tokens，并内置自适应思考、工具使用和结构化输出能力。它擅长高级编程、浏览器及计算机操作智能体、企业知识工作、金融与法律分析，以及需要持续判断和高可靠性的多步骤任务。",{"id":987,"groupId":95,"locale":934,"slug":96,"title":988,"pageType":17,"contentSource":24,"contentRef":989,"sort":15,"status":16,"icon":53,"description":990,"httpMethod":64,"hasToc":27},"01a06b45-c7f2-73b1-baa9-5e4b9a79c906","Claude Opus 4.7 经济版","01a06b6b-63fa-7981-9a18-a6a51a74e0d5","Claude Opus 4.7 是 Anthropic 正式发布的最强 Opus 模型，专为复杂推理、长周期智能体编程和高度自主的专业工作流而设计。它支持文本、图像和文件输入，并生成文本输出。该模型拥有 100 万 tokens 的上下文窗口、最多 128,000 个输出 tokens，并内置自适应思考、工具使用和结构化输出能力。它擅长高级编程、浏览器及计算机操作智能体、企业知识工作、金融与法律分析，以及需要持续判断和高可靠性的多步骤任务。",{"id":992,"groupId":102,"locale":934,"slug":103,"title":993,"pageType":17,"contentSource":24,"contentRef":994,"sort":15,"status":16,"icon":53,"description":995,"httpMethod":64,"hasToc":27},"01a06b45-c716-7d87-810b-6dd297265bbd","Claude Opus 4.6 经济版","01a06b6b-5ec1-7aa7-bde7-a5c4793a5176","Claude Opus 4.6 是 Anthropic 正式发布的最强 Opus 模型，专为复杂推理、长周期智能体编程和高度自主的专业工作流而设计。它支持文本、图像和文件输入，并生成文本输出。该模型拥有 100 万 tokens 的上下文窗口、最多 128,000 个输出 tokens，并内置自适应思考、工具使用和结构化输出能力。它擅长高级编程、浏览器及计算机操作智能体、企业知识工作、金融与法律分析，以及需要持续判断和高可靠性的多步骤任务。",{"id":997,"groupId":109,"locale":934,"slug":110,"title":998,"pageType":17,"contentSource":24,"contentRef":999,"sort":15,"status":16,"icon":53,"description":1000,"httpMethod":64,"hasToc":27},"01a06b45-c485-7759-a758-e15b586b547f","Claude Haiku 4.5 经济版","01a06b6b-4e51-7220-89e3-859c62c8aad4","Claude Haiku 4.5 是 Anthropic 速度最快、效率最高的模型。与更大的 Claude 模型相比，它以显著更低的成本和延迟提供接近前沿水平的智能。在推理、编程和计算机操作任务上，其性能可与 Claude Sonnet 4 相媲美，可为实时、高并发应用提供高级能力。\n\n作为首个支持扩展思考的 Haiku 模型，Haiku 4.5 提供可调节的推理深度、总结式或交错式思考，以及覆盖编程、Bash、网页搜索和计算机操作的工具辅助工作流。它在 SWE-bench Verified 上得分超过 73%，跻身全球领先编程模型之列，同时在子智能体编排、并行执行和大规模部署中保持高度响应。",{"id":1002,"groupId":116,"locale":934,"slug":10,"title":117,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":118,"description":1003,"hasToc":55,"children":1004},"01a06b45-3dfe-73e5-aea6-1e9664d1ebb0","平台现已正式接入 OpenAI 最新旗舰级 GPT-5.6 模型套件 API。GPT-5.6 放弃了单一架构设计，针对企业级开发与复杂生产管线推出了 Sol (旗舰推理版)、Terra (均衡主力版) 与 Luna (轻量极速版) 三款分级模型。其中 Sol 搭载了全新的 Max Reasoning 深度推理机制与 Parallel Subagents（并行子智能体）技术，能自主拆解并并行处理高难度的算法架构、网络安全推演与长文本分析；Terra 提供了媲美上一代旗舰的综合表现，同时推理成本降低 50%；Luna 则是高吞吐、低延迟任务的极速引擎",[1005,1010,1015,1020],{"id":1006,"groupId":123,"locale":934,"slug":124,"title":1007,"pageType":17,"contentSource":24,"contentRef":1008,"sort":15,"status":16,"icon":127,"description":1009,"httpMethod":64,"hasToc":27},"01a079c3-33e4-7470-a1f7-40fb292432f3","GPT 6 Astra 经济版","01a079c4-3948-7933-9f01-0c73e1f67040","GPT-6 Astra 是 OpenAI 推出的新一代前沿级（Frontier）旗舰大模型，被誉为具备 AGI 雏形的新一代智能底座。模型专为长流程（Long-horizon）复杂工作流与 Agent 自主协作打造，具备顶尖的深度逻辑推理、跨应用计算机操控（Computer Use）、高级代码构建与科学计算能力。基于强化的多步骤推演与符号世界建模机制，Astra 能在无 API 的软件环境中直接通过界面视觉交互自主执行任务、纠错并动态修正路径，在 ARC-AGI-3、软件工程与金融建模等前沿评估中展现出突破性表现，广泛适用于全流程软件开发、企业级 Agent 工作流编排、深度商业智能分析与高难度决策支撑。",{"id":1011,"groupId":131,"locale":934,"slug":132,"title":1012,"pageType":17,"contentSource":24,"contentRef":1013,"sort":15,"status":16,"icon":127,"description":1014,"httpMethod":64,"hasToc":27},"01a06b45-c8ce-7bb4-9b6c-8954e6e44a91","GPT 5.6 Terra 经济版","01a06b6b-6915-7347-951e-190958b8b552","GPT-5.6 Terra 是 OpenAI GPT-5.6 系列中的均衡型模型，定位介于旗舰级 Sol 和高性价比 Luna 之间。它非常适合日常编程、推理和智能体工作流，在通用生产环境中实现了质量、延迟和成本之间的良好平衡。",{"id":1016,"groupId":138,"locale":934,"slug":139,"title":1017,"pageType":17,"contentSource":24,"contentRef":1018,"sort":15,"status":16,"icon":142,"description":1019,"httpMethod":64,"hasToc":27},"01a06b45-ca07-7c62-aa7e-74d691f05ed7","GPT 5.6 Sol 经济版","01a06b6b-6e30-7f19-a5b4-8e91bee7f8f1","GPT-5.6 Sol 是 OpenAI GPT-5.6 系列的旗舰模型。它专为复杂推理、编程和智能体工作流而设计，尤其擅长多步骤问题解决、命令行辅助和高质量软件任务。当输出质量和可靠性比原始吞吐量更重要时，推荐使用该模型。",{"id":1021,"groupId":146,"locale":934,"slug":147,"title":1022,"pageType":17,"contentSource":24,"contentRef":1023,"sort":15,"status":16,"icon":127,"description":1024,"httpMethod":64,"hasToc":27},"01a06b45-cfad-7db5-9c3a-abace09e0c6f","GPT 5.5 经济版","01a06b6b-8efb-7908-b97f-931661bc7d9c","GPT-5.5 是 OpenAI 于 2026 年 4 月 23 日发布的前沿模型。它拥有超过 100 万 tokens 的上下文窗口，其中包括 922,000 个输入 tokens 和 128,000 个输出 tokens，并支持文本和图像输入。该模型在 SWE-bench Verified 上取得 88.7% 的成绩，在 MMLU 上取得 92.4% 的成绩；与 GPT-5.4 相比，其幻觉减少了 60%。它擅长智能体编程、计算机操作和深度研究，同时保持与 GPT-5.4 相当的单 token 延迟。",{"id":1026,"groupId":153,"locale":934,"slug":10,"title":1027,"pageType":32,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":155,"hasToc":27,"children":1028},"01a08417-4a1e-78a2-bc82-949118879017","Gemini",[1029,1033,1037],{"id":1030,"groupId":159,"locale":934,"slug":160,"title":161,"pageType":17,"contentSource":24,"contentRef":1031,"sort":15,"status":16,"icon":155,"description":1032,"httpMethod":64,"hasToc":27},"01a08431-1b14-7a02-84da-969341e71849","01a08017-4ccb-7a69-b5ac-0d8ba2774d92","Gemini 3.7 Flash 是谷歌推出的新一代主力高性价比全模态 AI 大模型。模型具备原生多模态理解与深度推理能力，针对长流程 Agent 工作流、多步骤复杂工具调用与代码工程进行了算法级优化。凭借更低延迟、低至 $0.75\u002F1M 输入 Token 的超高性价比与自适应策略调整机制，Gemini 3.7 Flash 广泛适用于大规模企业级 Agent 部署、高并发实时交互、全自动化软件开发与多模态数据处理。",{"id":1034,"groupId":166,"locale":934,"slug":167,"title":168,"pageType":17,"contentSource":24,"contentRef":1035,"sort":15,"status":16,"icon":155,"description":1036,"httpMethod":64,"hasToc":27},"01a08434-834a-7f94-b6ce-5a5ad5831f04","01a06b6a-c555-7aed-9cab-b7077e980166","Gemini 3.6 Flash 是 Google 于 2026 年 7 月推出的新一代高效轻量级主力大模型。专为智能体（Agent）工作流、复杂代码构建与多模态任务打造，支持 100 万 Token 上下文输入与 6.4 万 Token 输出限制。相较于 3.5 Flash，其输出 Token 消耗降低约 17%，工具调用与推理步数更加精简，大幅提升了 agent 任务的执行效率与性价比。原生支持文本、图像、视频、音频及 PDF 解析，具备出色的电脑操作（Computer Use）与多工具协同能力。",{"id":1038,"groupId":173,"locale":934,"slug":174,"title":175,"pageType":17,"contentSource":24,"contentRef":1039,"sort":15,"status":16,"icon":155,"description":1040,"httpMethod":64,"hasToc":27},"01a08435-9785-7909-8526-264d11b17596","01a06b6a-aa87-7de3-ab00-80685a341554","Gemini 3.5 Flash 是 Google 推出的高效轻量级多模态主力大模型。专为高并发、低延迟的智能体（Agent）工作流、代码构建与多模态理解打造，支持 100 万 Token 长上下文窗口。原生支持文本、图像、视频、音频及文档解析，具备极致的推理速度与极高的性价比，在工具调用（Tool Use）、逻辑推理与多语言处理上表现优异，广泛适用于企业级 API 集成与实时交互场景。",{"id":1042,"groupId":180,"locale":934,"slug":10,"title":181,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":182,"description":1043,"hasToc":55,"children":1044},"01a06b45-379c-7be0-9348-125ed3196c2f","平台现已正式接入智谱（Z.ai \u002F 智谱 AI）新一代旗舰级 GLM-5.3 大模型生产级 API。GLM-5.3 延续了高性能混合专家（MoE）架构，通过数十倍长程任务环境的极致后训练（Post-Training Scaling），实现了智能上限与有效智能效率的质变。模型在复杂代码编写、长文本逻辑推演、终端自动化指令集执行（Terminal-Bench）及防御性网络安全（CyberGym）等前沿任务中达到了顶尖水平。",[1045,1049,1053],{"id":1046,"groupId":201,"locale":934,"slug":202,"title":203,"pageType":17,"contentSource":24,"contentRef":1047,"sort":15,"status":16,"icon":182,"description":1048,"httpMethod":64,"hasToc":27},"01a06b45-a8e9-7771-b13e-69fa6ebf8a5a","01a06b6a-a583-7228-8ac5-5511a26cce08","GLM 5.2 是智谱（Z.ai）推出的大规模推理模型。它支持文本输入和输出，拥有 100 万 Token 的上下文窗口，适用于长周期智能体工作流、项目级软件工程和复杂的多步骤自动化任务。\n\n该模型支持 `high` 和 `xhigh` 两种推理强度，其中 `xhigh` 对应最高推理级别。它尤其擅长长时间运行任务中的编程与工具使用，能够持续保持工程上下文并始终遵循相关规范，在单个任务中完成从需求分析到多平台部署的完整开发流程。",{"id":1050,"groupId":187,"locale":934,"slug":188,"title":189,"pageType":17,"contentSource":24,"contentRef":1051,"sort":15,"status":16,"icon":182,"description":1052,"httpMethod":64,"hasToc":27},"01a06b45-ab97-7ac4-8b44-11f891c2a577","01a06b6a-b49e-73cc-9346-466bfc331b9a","GLM 5.3 Flash 是智谱 AI 推出的新一代极速轻量级主力大模型。专为高并发、低延迟的 Agent 智能体工作流、代码构建与多模态任务打造。模型原生支持长上下文输入，推理吞吐量大幅提升，兼具极高的性价比。在工具调用（Function Calling）、指令遵循、逻辑推理及多语言理解方面表现突出，广泛适用于企业级 API 高效接入、实时交互与自动化工作流场景。",{"id":1054,"groupId":194,"locale":934,"slug":195,"title":196,"pageType":17,"contentSource":24,"contentRef":1055,"sort":15,"status":16,"icon":182,"description":1056,"httpMethod":64,"hasToc":27},"01a06b45-b2cc-7b6c-bb9c-d7dbab19ce44","01a06b6a-e037-724e-a5a3-e5a6e0b6e950","GLM 5.3 是智谱（Z.ai）推出的大规模推理模型。它支持文本输入和输出，拥有 100 万 Token 的上下文窗口，适用于长周期智能体工作流、项目级软件工程和复杂的多步骤自动化任务。\n\n该模型支持 `high` 和 `xhigh` 两种推理强度，其中 `xhigh` 对应最高推理级别。它尤其擅长长时间运行任务中的编程与工具使用，能够持续保持工程上下文并始终遵循相关规范，在单个任务中完成从需求分析到多平台部署的完整开发流程。",{"id":1058,"groupId":208,"locale":934,"slug":10,"title":209,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":210,"description":1059,"hasToc":55,"children":1060},"01a06b45-3cc9-7d48-ae23-9c301c77e723","平台现已全面接入深度求索（DeepSeek）最新旗舰级 DeepSeek V4 模型套件生产级 API。作为全球开源 AI 领域的标杆，DeepSeek V4 套件包含 DeepSeek-V4-Pro (旗舰深度推理版) 与 DeepSeek-V4-Flash (高吞吐极速版) 两款核心模型。V4 架构在复杂逻辑推演、长上下文检索与长链条 Agent 自动编排上实现了质的飞跃，针对代码重构（Codex）、终端指令控制与自动化工作流进行了深度优化。",[1061,1065],{"id":1062,"groupId":215,"locale":934,"slug":216,"title":217,"pageType":17,"contentSource":24,"contentRef":1063,"sort":15,"status":16,"icon":210,"description":1064,"httpMethod":64,"hasToc":27},"01a06b45-a809-7284-9c5e-7da99c5f3db3","01a06b6a-a07a-7363-8bb4-ed304b48c00f","DeepSeek-V4-Pro-0813 是 2026 年 8 月 13 日发布的旗舰级 MoE 大模型，专为复杂逻辑推理、代码构建与 Agent 智能体协同打造的 GA 正式版。1.6T 参数（49B 激活），内置 DSpark 投机解码，推理吞吐显著提升。兼容 OpenAI\u002FAnthropic API，深度集成 DeepSeek Harness，在 Terminal Bench 表现优异。",{"id":1066,"groupId":222,"locale":934,"slug":223,"title":224,"pageType":17,"contentSource":24,"contentRef":1067,"sort":15,"status":16,"icon":210,"description":1068,"httpMethod":64,"hasToc":27},"01a06b45-ced1-7685-8392-89e6a21403c4","01a06b6b-8a19-72e1-9224-b53199b234e2","DeepSeek V4 Flash 是 DeepSeek 推出的效率优化型混合专家（MoE）模型，拥有 2840 亿总参数，每个 token 激活 130 亿参数，并提供 100 万 tokens 的上下文窗口。它专为快速推理和高吞吐量任务而设计，在保持出色成本效率的同时，提供强大的推理与编程能力。\n\n其混合注意力架构可实现高效的长上下文处理。该模型支持 **high** 和 **xhigh** 推理级别，其中 **xhigh** 代表最高推理强度。它非常适合对响应速度、可扩展性和成本效率有较高要求的编程助手、对话系统及智能体工作流。",{"id":1070,"groupId":264,"locale":934,"slug":10,"title":265,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":266,"description":1071,"hasToc":55,"children":1072},"01a06b45-3ed5-726f-99cc-a9972dc8e98e","平台已接入 MiniMax M3 与 M2.7 旗舰大模型 API。M3 首创 MSA 稀疏注意力，支持 1M 超长上下文、原生全模态与顶尖代码重构；M2.7 基于自我演进架构，专精 Agent Teams 协作与复杂办公文档处理。极低延迟、支持 Prompt Cache，为企业级 Agent 与 AI 编程提供生产级基础设施。",[1073,1077],{"id":1074,"groupId":271,"locale":934,"slug":272,"title":273,"pageType":17,"contentSource":24,"contentRef":1075,"sort":15,"status":16,"icon":266,"description":1076,"httpMethod":64,"hasToc":27},"01a06b45-afe9-7414-b2b5-cd15998af915","01a06b6a-cf85-7bc2-8ad6-59f6d617b519","MiniMax-M3 是 MiniMax 最新的 M 系列多模态基础模型，专为智能体推理、工具使用、编程和长上下文任务而构建。它支持文本、图像和视频输入，并生成文本输出，同时提供 100 万 tokens 的上下文窗口、扩展思考、函数调用和结构化输出能力。凭借在长周期智能体工作流、软件开发、多模态理解和长回复生成方面的强大能力，MiniMax-M3 非常适合自主智能体、编程助手、文档与视频分析，以及以具有竞争力的成本处理超大上下文的生产级应用。",{"id":1078,"groupId":278,"locale":934,"slug":279,"title":280,"pageType":17,"contentSource":24,"contentRef":1079,"sort":15,"status":16,"icon":266,"description":1080,"httpMethod":64,"hasToc":27},"01a06b45-b0c4-7947-a39b-b06b5ea12725","01a06b6a-d4fc-7169-ac16-d541f20957ff","MiniMax-M2.7 是面向现实世界自主生产力和持续改进而构建的新一代大语言模型。凭借先进的智能体能力和多智能体协作，它可以在动态环境中规划、执行、评估和优化复杂任务，同时积极参与自身能力的演进。\n\nM2.7 针对生产级工作流进行了优化，擅长实时调试、根因分析、金融建模，以及跨 Word、Excel 和 PowerPoint 的端到端文档创建。它在 SWE-Pro 上取得 56.2%、在 Terminal Bench 2 上取得 57.0%，并在 GDPval-AA 上获得 1495 ELO 评分，为现实数字工作流中的多智能体系统树立了新基准。",{"id":1082,"groupId":285,"locale":934,"slug":10,"title":286,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":287,"description":1083,"hasToc":55,"children":1084},"01a06b45-3fb5-7b22-befe-68704ed994b1","平台现已正式接入月之暗面（Moonshot AI）最新旗舰级 Kimi K3 及 K2 系列模型生产级 API。模型不仅具备极强的超长上下文无损处理能力与 Context Caching（上下文缓存）降本提速机制，更在深度推理（K2\u002FK3 Thinking）、复杂代码工程、数据分析与自动化 Agent 编排中展现出顶尖性能。",[1085],{"id":1086,"groupId":292,"locale":934,"slug":293,"title":294,"pageType":17,"contentSource":24,"contentRef":1087,"sort":15,"status":16,"icon":287,"description":1088,"httpMethod":64,"hasToc":27},"01a06b45-b1a4-78e0-b406-69b7b6380eee","01a06b6a-dade-7d1f-8096-df4f05488b23","Kimi K3 是 Kimi 迄今能力最强的旗舰模型。它拥有 2.8 万亿参数，基于 Kimi Delta Attention（KDA）混合线性注意力架构和 Attention Residuals 技术构建，原生支持视觉理解，并提供 100 万 tokens 的上下文窗口。\n\n作为全球首个接近 3 万亿参数规模的开源模型，Kimi K3 专为前沿 AI 应用而设计，包括长周期编程、知识工作和推理任务。",{"id":1090,"groupId":229,"locale":934,"slug":10,"title":230,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":231,"description":1091,"hasToc":55,"children":1092},"01a06b45-4530-74e4-8b03-bd66af61ccf0","平台现已全系接入阿里云（Alibaba Cloud）最新旗舰级 Qwen 3.8 及 Qwen3 系列模型套件生产级 API。作为全球开源与商业 AI 大模型领域的标杆，Qwen 家族提供了包含 Qwen3.8-Max (2.4T 顶级推理与编程版)、Qwen3.7-Plus (高性价比全能版)、Qwen3.8-27B (极速端侧\u002F私有化版) 以及 Qwen-Image-3.0 (高保真图像生成版) 的全方位矩阵。全新一代 Qwen 在复杂数学逻辑推演、深度代码重构（Coding）、Agent 世界模型环境交互（AgentWorld）与多语言理解上实现了全面突破。",[1093,1097,1101,1105],{"id":1094,"groupId":236,"locale":934,"slug":237,"title":238,"pageType":17,"contentSource":24,"contentRef":1095,"sort":15,"status":16,"icon":231,"description":1096,"httpMethod":64,"hasToc":27},"01a06b45-ac73-7795-9f8c-c520092f45c3","01a06b6a-b99f-70aa-8ee8-b269e7c68954","Qwen3.8-Max 是阿里巴巴 Qwen3.8 系列的旗舰模型，专为基于文本的智能体工作流而设计。它擅长编程、调试、办公自动化、生产力任务、工具使用和长周期自主执行。该模型拥有 100 万 tokens 的上下文窗口，并支持最多 64K tokens 的输出，非常适合处理大型文档、代码仓库级编程、多步骤规划、结构化内容生成，以及需要跨数百甚至数千个步骤持续推理的复杂工作流。",{"id":1098,"groupId":243,"locale":934,"slug":244,"title":245,"pageType":17,"contentSource":24,"contentRef":1099,"sort":15,"status":16,"icon":231,"description":1100,"httpMethod":64,"hasToc":27},"01a06b45-b3a5-78c4-9aba-5b478811f03f","01a06b6a-e555-7600-b807-c45fcfc42c13","Qwen3.7-Max 是阿里巴巴 Qwen3.7 系列的旗舰模型，专为基于文本的智能体工作流而设计。它擅长编程、调试、办公自动化、生产力任务、工具使用和长周期自主执行。该模型拥有 100 万 tokens 的上下文窗口，并支持最多 64K tokens 的输出，非常适合处理大型文档、代码仓库级编程、多步骤规划、结构化内容生成，以及需要跨数百甚至数千个步骤持续推理的复杂工作流。",{"id":1102,"groupId":250,"locale":934,"slug":251,"title":252,"pageType":17,"contentSource":24,"contentRef":1103,"sort":15,"status":16,"icon":231,"description":1104,"httpMethod":64,"hasToc":27},"01a06b45-b47f-7ab5-901a-e0b4308a24a9","01a06b6a-ea7e-7232-8b63-f9295bde6db8","Qwen3.7-Plus 是阿里巴巴 Qwen3.7 系列中的高性价比模型，支持文本和图像输入，并生成文本输出。它将增强的视觉语言能力与面向编程、工具使用和生产力工作流的全栈智能体能力相结合。其核心优势是多模态交互式智能体功能：它可以理解现实场景、解读屏幕内容、与图形界面交互、根据视觉参考生成代码，并自主完成移动应用的端到端操作。",{"id":1106,"groupId":257,"locale":934,"slug":258,"title":259,"pageType":17,"contentSource":24,"contentRef":1107,"sort":15,"status":16,"icon":231,"description":1108,"httpMethod":64,"hasToc":27},"01a06b45-aa9d-7649-b04d-8810b9f12b8d","01a06b6a-af9d-7272-8ab2-a1a08f0b541e","Qwen 3.5 Omni Plus 是阿里云推出的新一代端到端全模态（Omni）旗舰大模型。模型采用原生多模态融合架构，支持文本、图像、音频与视频的多模态统一输入输出与实时双工交互。具备超低延迟流式推理、高逼真情感语音合成、音画联合推理与复杂运镜\u002F视觉解析能力，在 Agent 工具调用与跨模态理解上表现卓越，广泛适用于实时语音助手、智能客服、视频交互与 AI 数字人等场景。",{"id":1110,"groupId":299,"locale":934,"slug":10,"title":1111,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":301,"description":1112,"hasToc":55,"children":1113},"01a06b45-3419-77a4-8e2f-069c68203e01","腾讯 Hy","腾讯Hy系列是腾讯公司（Tencent）自主研发的旗舰级 Hunyuan 3.0 \u002F Hy3\u002F Hy 4 大语言模型生产级 API。Hy3 采用先进的 295B 稀疏 MoE 架构（单 Token 推理仅激活 21B 参数），实现了卓越的推理精度与极致的算力成本控制。模型独创“快慢结合思考”机制，支持最高 256K 无损上下文理解，在复杂代码重构（SWE-Bench）、多步长路线智能体调度（QClaw \u002F WorkBuddy）、深度科学推演与中文复杂语义检索上展现出业内顶尖的性能表现。",[1114,1119],{"id":1115,"groupId":306,"locale":934,"slug":307,"title":1116,"pageType":17,"contentSource":24,"contentRef":1117,"sort":15,"status":16,"icon":301,"description":1118,"httpMethod":64,"hasToc":27},"01a06b45-af0e-7572-918d-9afd6022801d","腾讯 Hy4 Preview","01a06b6a-ca77-72d6-877b-6f18aa6fbfe0","腾讯混元 Hy4 preview（Tencent Hy4 preview）是腾讯推出的新一代 770B MoE 旗舰级开源大模型。模型拥有 7700 亿总参数与 49 亿激活参数，支持 100 万 Token（1M）超长上下文。架构采用带 IndexCache 的 Gated DSA 稀疏注意力机制与 iHC 残差连接，并内置 10B 参数的 MTP 投机采样层。",{"id":1120,"groupId":313,"locale":934,"slug":314,"title":1121,"pageType":17,"contentSource":24,"contentRef":1122,"sort":15,"status":16,"icon":301,"description":1123,"httpMethod":64,"hasToc":27},"01a06b45-ad5a-7aa1-aabe-50db10651b04","腾讯 Hy3","01a06b6a-c033-76b2-8a09-eebda6f62155","Tencent Hy3（腾讯混元 3）是腾讯推出的新一代旗舰级 MoE 架构大模型。模型拥有 2950 亿总参数与 210 亿激活参数，采用多 Token 预测（MTP）架构，专为 Agent 智能体工作流、复杂代码构建、长上下文处理与逻辑推理打造。",{"id":1125,"groupId":320,"locale":934,"slug":10,"title":321,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":322,"description":1126,"hasToc":55,"children":1127},"01a06b45-34f2-78f9-a5c2-411fc225040b","Xiaomi MiMo（小米 MiMo）是小米集团自主研发的新一代 AI 大模型系列，作为“人车家全生态”战略的核心 AI 底座。模型具备原生全模态理解（文本、图像、语音与视频）、复杂逻辑推理、代码生成与智能体（Agent）自主协同能力，最高支持 100 万 Token 长上下文与超高吞吐量流式推理。",[1128],{"id":1129,"groupId":327,"locale":934,"slug":328,"title":329,"pageType":17,"contentSource":24,"contentRef":1130,"sort":15,"status":16,"icon":322,"description":1131,"httpMethod":64,"hasToc":27},"01a08015-0ed1-74a7-8f3f-7292f4e3be43","01a08017-5298-74cf-99cb-a952f891cd45","MiMo V2.5 Pro（小米 MiMo V2.5 Pro）是小米推出的旗舰级开源 MoE 智能体大模型。模型拥有 1.02 万亿总参数与 420 亿激活参数，支持 100 万 Token（1M）超长上下文。架构采用滑动窗口与全局混合注意力机制（Hybrid Attention），并内置 MTP 多 Token 预测层。模型主打极致的 Token 效率与长程复杂任务执行能力，可支撑千次以上的 Tool Call 自主推理与自我纠错，广泛适用于软件工程全流程开发、复杂 Agent 工作流编排、硬件 EDA 仿真与人车家全场景生态。",{"id":1133,"groupId":334,"locale":934,"slug":10,"title":1134,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"topTab":17,"children":1135},"01a06b0c-6499-7302-a795-097c7a9c9517","图片模型",[1136,1152,1163,1175,1183,1196,1209],{"id":1137,"groupId":339,"locale":934,"slug":10,"title":340,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":127,"description":1138,"hasToc":55,"children":1139},"01a08513-6b45-77b9-8579-e39664712e45","GPT Images 2.5 是 OpenAI 推出的新一代旗舰级商业视觉生成与精准图像编辑大模型。模型包含 Standard（擅长真实摄影与插画）与 Ultra（超高清晰度与极佳文本渲染）两个版本，支持最高 10 张参考图精准锁定人物外观、产品细节、设计风格与 Logo 结构，实现精准的“多图一致性”（Multi-image Identity Consistency）。结合草图引导（Sketch）与模板合成（Templates），用户可以精准控制画面构图并实现自动化商品图替换。凭借出色的多文字排版控制与原生 4K 商业画质，GPT Images 2.5 全面赋能跨境电商视觉营销、品牌海报设计、UI\u002FUX 原型图与高转化率广告内容创作。",[1140,1144,1148],{"id":1141,"groupId":345,"locale":934,"slug":346,"title":1142,"pageType":17,"contentSource":24,"contentRef":1143,"sort":15,"status":16,"icon":118,"description":1138,"httpMethod":64,"hasToc":27},"01a08513-ae60-7248-8c1a-62556f8a4f62","GPT Image 2.5 经济版","01a08512-9a9b-728e-a80e-ec20c449066b",{"id":1145,"groupId":352,"locale":934,"slug":353,"title":1146,"pageType":17,"contentSource":24,"contentRef":1147,"sort":15,"status":16,"icon":118,"httpMethod":64,"hasToc":27},"01a09016-4bc2-752a-8ec7-0c8b1aca105b","GPT Image 2.5 Flare 经济版","01a09015-937c-7be1-8151-0e1d6ee2e6d6",{"id":1149,"groupId":358,"locale":934,"slug":359,"title":1150,"pageType":17,"contentSource":24,"contentRef":1151,"sort":15,"status":16,"icon":118,"httpMethod":64,"hasToc":27},"01a09016-4d6b-7778-9936-2e6224b42827","GPT Image 2.5 Sunburst 经济版","01a09015-9da0-77b9-8636-9f73d5abf2f4",{"id":1153,"groupId":364,"locale":934,"slug":10,"title":365,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":118,"description":1154,"hasToc":55,"children":1155},"01a06b45-3b12-7e0f-ab9e-0107ab35ed80","GPT Image 2 首次将原生“思维链推理”（Thinking Mode）深度融入图像生成全流程。模型在生成前会自动规划空间构图、检索实时真实参照，并在直出 2K 极清视觉画面后进行自我校验优化。除了拥有极致的光影质感与真实摄影细节，GPT Image 2 彻底解决了长段文本渲染与复杂 UI 设计图的生成痛点，并支持无缝的“对话式局部重绘”（Targeted Inpainting）与多镜头下连贯的人物角色塑造。",[1156,1160],{"id":1157,"groupId":370,"locale":934,"slug":371,"title":365,"pageType":17,"contentSource":24,"contentRef":1158,"sort":15,"status":16,"icon":127,"description":1159,"httpMethod":64,"hasToc":27},"01a06b45-979f-7aed-822b-d5403c0dd829","01a06b6a-483c-710c-bccb-83f0b07a267a","OpenAI 的 GPT Image 2 原生图像模型可根据自然语言提示词生成高质量图像。它提供开箱即用的 REST 推理 API，性能出色、无冷启动且价格实惠。",{"id":1161,"groupId":376,"locale":934,"slug":377,"title":378,"pageType":17,"contentSource":24,"contentRef":1162,"sort":15,"status":16,"icon":127,"description":1159,"httpMethod":64,"hasToc":27},"01a06b45-9879-7fea-a7e8-c0b94218528e","01a06b6a-4de5-70e0-b408-382cff1a6ec2",{"id":1164,"groupId":382,"locale":934,"slug":10,"title":383,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":384,"description":1165,"hasToc":55,"children":1166},"01a06b45-3342-7d93-8082-04eba17cf371","字节跳动（ByteDance）最新一代旗舰图像生成大模型 Seedream 5.0 Pro 现已全面提供生产级 API 接入。作为视觉生成领域的工业级标杆，5.0 Pro 版本在光学材质渲染、真实光影细节以及构图控制力上实现了全面飞跃，原生支持高达 4K 的商业级高清图像输出。模型引入了精准的 HEX 色彩代码控制，可严格遵循品牌视觉规范，并在复杂画面文字排版（包含多语言平面设计与海报布局）及多图角色\u002F主体一致性上展现出顶尖的商业化落地能力。",[1167,1171],{"id":1168,"groupId":389,"locale":934,"slug":390,"title":391,"pageType":17,"contentSource":24,"contentRef":1169,"sort":15,"status":16,"icon":384,"description":1170,"httpMethod":64,"hasToc":27},"01a06b45-a04d-796d-9f53-4f8deefd8c1c","01a06b6a-77a4-7de9-8cb3-4d19737e1c8d","Seedream 5.0 Pro 文生图 辣味模式（高饱和\u002F高动态版）是无审查版字节跳动推出的一款专注于极致视觉表现力的专业级文生视频大模型。作为 5.0 Pro 系列的超强衍生版，它在保持专业级细节与可控性的基础上，针对画面的色彩饱和度、动态范围（HDR）、光影对比度及视觉张力进行了深度优化。",{"id":1172,"groupId":396,"locale":934,"slug":397,"title":398,"pageType":17,"contentSource":24,"contentRef":1173,"sort":15,"status":16,"icon":384,"description":1174,"httpMethod":64,"hasToc":27},"01a06b45-a208-7486-b01d-517b05f9bbe0","01a06b6a-81f1-7d6b-b81a-602b30176b86","Seedream 5.0 Pro Edit Spicy（高饱和度\u002F高动态范围版）是由 Seedream AI 推出的专业级 AI 图像编辑模型变体，专为营造极致视觉冲击力而打造。该版本在继承 5.0 Pro 系列专业级细节表现与精准控制力的基础上，进行了针对性微调，旨在实现极高的色彩饱和度、宽广的动态范围（HDR）、强烈的明暗光影对比以及极具张力的视觉效果。",{"id":1176,"groupId":403,"locale":934,"slug":10,"title":404,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":384,"description":1177,"hasToc":55,"children":1178},"01a06b45-387b-7d75-880b-83b31a79adf1","Seedream 5.0 Lite 将字节跳动（ByteDance）前沿的图像生成技术带入对生成效率与成本控制有极高要求的生产场景中。作为 Seedream 5.0 旗舰矩阵的高效轻量化版本，Lite 模型在大幅提升渲染速度、显著降低单图生成成本的同时，完美继承了 5.0 系列的核心能力——包括高精度的文字排版、多语言海报布局以及出色的画面构图控制。",[1179],{"id":1180,"groupId":409,"locale":934,"slug":410,"title":404,"pageType":17,"contentSource":24,"contentRef":1181,"sort":15,"status":16,"icon":384,"description":1182,"httpMethod":64,"hasToc":27},"01a06b45-9f6e-7a64-8462-4046e914ddb0","01a06b6a-7282-74ba-b4df-a1fe19878373","Seedream 5.0 是一款先进的文生图模型，具备增强的文字排版能力，可为海报和品牌视觉设计清晰呈现文字，同时拥有出色的提示词遵循能力，并支持最高 4K 分辨率。提供开箱即用的 REST 推理 API，性能卓越、无冷启动，且价格实惠。",{"id":1184,"groupId":415,"locale":934,"slug":10,"title":416,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":417,"description":1185,"hasToc":55,"children":1186},"01a06b45-3bea-70c4-aab7-3af320020bfc","Nano Banana 摒弃了传统 AI 绘图的局限，深度融合谷歌真实世界知识库与多轮对话编辑能力。模型原生支持 1K、2K 及高达 4K (4096×2304) 的极清画面渲染，具备行业顶尖的商业级多语言文字排版与设计海报生成能力。在复杂的图生图与迭代编辑管线中，最高支持同时输入多达 14 张参考图像，确保角色外观、服装风格与特定主体在多场景创作中具备极高的一致性。",[1187,1191],{"id":1188,"groupId":422,"locale":934,"slug":423,"title":416,"pageType":17,"contentSource":24,"contentRef":1189,"sort":15,"status":16,"icon":417,"description":1190,"httpMethod":64,"hasToc":27},"01a06b45-b81a-74d6-b52b-0f79a346ea8b","01a06b6a-ff84-733f-9f21-3e5b0f647e41","Google Nano Banana Pro（Gemini 3.0 Pro Image）支持图像编辑，并可输出 4K 分辨率的结果。它提供开箱即用的 REST 推理 API，性能出色、无冷启动且价格实惠。",{"id":1192,"groupId":428,"locale":934,"slug":429,"title":1193,"pageType":17,"contentSource":24,"contentRef":1194,"sort":15,"status":16,"icon":417,"description":1195,"httpMethod":64,"hasToc":27},"01a06b45-b8f4-7d38-a254-87be82fbac58","Nano Banana Pro 文生图","01a06b6b-04a8-7927-8cea-48066a3b0cd2","Google Nano Banana Pro（Gemini 3.0 Pro Image）支持文本生图，并可输出 4K 分辨率的结果。它提供开箱即用的 REST 推理 API，性能出色、无冷启动且价格实惠。",{"id":1197,"groupId":435,"locale":934,"slug":10,"title":436,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":417,"description":1198,"hasToc":55,"children":1199},"01a06b45-4607-75ef-856a-ae42c591300c","平台现已正式接入由 Google 旗舰级 Gemini 3.1 Flash Image 模型驱动的 Nano Banana 2 生产级图像生成大模型 API。作为结合了 Pro 级视觉质感与 Flash 级极致生成速度的第二代旗舰模型，Nano Banana 2 突破性支持原生 4K (4096×2304) 高清画面直接渲染，无需二次放大。模型深度融合了谷歌实时真实世界知识（Real-World Knowledge）与搜索增强能力，具备业内顶尖的商业级多语言文字排版与复杂的图表设计能力。",[1200,1205],{"id":1201,"groupId":441,"locale":934,"slug":442,"title":1202,"pageType":17,"contentSource":24,"contentRef":1203,"sort":15,"status":16,"icon":417,"description":1204,"httpMethod":64,"hasToc":27},"01a06b45-b735-707c-b532-62ecd5735394","Nano Banana 2 文生图","01a06b6a-fa66-7de8-acef-83b42aff66c9","Nano Banana 2 Text-to-Image（Gemini 3.1 Flash Image）是 Google 新一代图像编辑与生成 AI 模型，让视觉创作像使用文字描述一样简单直观。该模型基于 Google 前沿的计算机视觉和生成式 AI 技术，结合精准控制、创作灵活性与深层语义理解，可提供专业级的图像编辑与生成能力。",{"id":1206,"groupId":448,"locale":934,"slug":449,"title":436,"pageType":17,"contentSource":24,"contentRef":1207,"sort":15,"status":16,"icon":417,"description":1208,"httpMethod":64,"hasToc":27},"01a06b45-bba8-7798-a8c0-9b3c576b5645","01a06b6b-1559-7fcb-9fd1-3df331995a6b","Nano Banana 2（Gemini 3.1 Flash Image）是 Google 新一代图像编辑与生成 AI 模型，让视觉创作像使用文字描述一样简单直观。该模型基于 Google 前沿的计算机视觉和生成式 AI 技术，结合精准控制、创作灵活性与深层语义理解，可提供专业级的图像编辑与生成能力。",{"id":1210,"groupId":454,"locale":934,"slug":10,"title":455,"pageType":32,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":231,"description":1091,"hasToc":27,"children":1211},"01a0ad89-e2fa-7493-a8e9-cc82f00f0d0f",[1212,1216],{"id":1213,"groupId":459,"locale":934,"slug":460,"title":461,"pageType":17,"contentSource":24,"contentRef":1214,"sort":15,"status":16,"icon":231,"description":1215,"httpMethod":64,"hasToc":27},"01a079c3-374e-72d7-bfaf-519dcdc2ad28","01a06c19-16cc-75ff-8bb5-86b387ce2c2c","Qwen Image 3.0（通义千问 Qwen Image 3.0）是阿里巴巴推出的新一代开源 AI 图像生成与编辑大模型。基于原生 Diffusion Transformer（DiT）架构，模型同时支持高精度文生图（T2I）与图生图\u002F指令编辑（I2I）。主打“富文本渲染、真实细节与深厚知识”，支持低至 10px 的微小文字清晰渲染与 12 种语言原生排版，可一键生成复杂的图文排版、报纸版面与 UI 界面。模型具备微米级材质细节还原、高通透物理光影与多图编辑能力，开放权重支持企业自建私有化部署，广泛适用于品牌海报设计、UI\u002FUX 原型图、跨境电商营销与多语言视觉创作。",{"id":1217,"groupId":466,"locale":934,"slug":467,"title":468,"pageType":17,"contentSource":24,"contentRef":1218,"sort":15,"status":16,"icon":231,"description":1219,"httpMethod":64,"hasToc":27},"01a079c3-38de-7c0e-b527-9cdb465c383b","01a06c19-1bbe-7311-b29e-c1a69eb32a6e","Qwen Image 3.0 Pro（通义千问 Qwen Image 3.0 Pro）是阿里巴巴推出的新一代旗舰级 AI 图像生成大模型。基于 Diffusion Transformer 架构，模型主打“高信息密度与实用级生产力”，最高支持 4,500 Token 超长 Prompt 输入，可一键生成报纸版面、多格分镜、学术图表与复杂 UI 界面。模型具备微米级细节还原与极佳的文字渲染能力，支持低至 10px 的微小文字清晰显示与 12 种语言原生排版；同时具备多图指令级编辑与深厚的世界知识理解，广泛适用于品牌广告设计、多语言营销海报、UI 原型探索、短剧分镜与电商高精视觉创作。",{"id":1221,"groupId":473,"locale":934,"slug":10,"title":1222,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"topTab":17,"children":1223},"01a06b0c-657b-73fb-96f0-d0a9ee572384","视频模型",[1224,1242,1274,1291,1314,1326,1344,1365,1383],{"id":1225,"groupId":478,"locale":934,"slug":10,"title":479,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":384,"description":1226,"hasToc":55,"children":1227},"01a06b45-3958-7498-984e-f3db047ffa26","字节跳动（ByteDance）最新一代旗舰视频生成大模型 Seedance 2.5 现已全面上线。模型支持通过文本、单张图片或多达 50 组多模态参考（包含图片、视频与音频），单次生成长达 30 秒的完整商业级视频。Seedance 2.5 突破性地集成了原生音画同步生成与画面内多语言字幕\u002F文本渲染，并在物理渲染仿真与主体长帧一致性上实现了全新跃迁，完美保持复杂长镜头中的角色与场景连贯性。",[1228,1232,1237],{"id":1229,"groupId":484,"locale":934,"slug":485,"title":479,"pageType":17,"contentSource":24,"contentRef":1230,"sort":15,"status":16,"icon":384,"description":1231,"httpMethod":64,"hasToc":27},"01a06b45-cc04-7acb-8718-072baac252e9","01a06b6b-7a5b-7d90-a5e2-2afcdd4d504c","Seedance 2.5 是一款基于参考图像、视频和音频进行多模态视频生成的工具，并支持视频编辑与扩展功能。",{"id":1233,"groupId":490,"locale":934,"slug":491,"title":1234,"pageType":17,"contentSource":24,"contentRef":1235,"sort":15,"status":16,"icon":384,"description":1236,"httpMethod":64,"hasToc":27},"01a06b45-ccdf-7ec7-9441-f16779f383e6","Seedance 2.5 文生视频","01a06b6b-7fb8-7c2d-8b42-fd19c0616a2f","Seedance 2.5 Text-to-Video 是一款基于参考文字进行多模态视频生成的工具，并支持视频编辑与扩展功能。",{"id":1238,"groupId":497,"locale":934,"slug":498,"title":1239,"pageType":17,"contentSource":24,"contentRef":1240,"sort":15,"status":16,"icon":384,"description":1241,"httpMethod":64,"hasToc":27},"01a06b45-cdbd-709e-bee5-f4d082b34258","Seedance 2.5 图生视频","01a06b6b-84b8-7f47-ad2c-dc7c6c719eca","Seedance 2.5 Image-to-Video 是一款基于参考文字进行多模态视频生成的工具，并支持视频编辑与扩展功能。",{"id":1243,"groupId":504,"locale":934,"slug":10,"title":505,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":384,"description":1244,"hasToc":55,"children":1245},"01a06b45-35cc-7dc4-9c14-b760070d9461","平台现已全面提供字节跳动（ByteDance）Seedance 2.0 旗舰视频模型的生产级 API 接入。Seedance 2.0 采用统一的多模态音视频联合生成架构，原生支持文本、图像、视频与音频四模态联合输入。模型搭载行业领先的“Universal Reference”（通用参考）系统，能够在多镜头切分与复杂长镜头中精准锁定画面构图、主体角色外观与运镜轨迹，真正实现导演级的高精度内容控制。现已重磅支持原生 4K 超高清渲染输出，通过统一 API 提供透明的按秒计费模式，配合企业级 SLA 正常运行时间保障，助力开发者与影视团队秒级调用，无缝构建商业级视听创作管线。",[1246,1250,1255,1260,1265,1270],{"id":1247,"groupId":510,"locale":934,"slug":511,"title":512,"pageType":17,"contentSource":24,"contentRef":1248,"sort":15,"status":16,"icon":384,"description":1249,"httpMethod":64,"hasToc":27},"01a06b45-be3c-7609-9016-57880bf6a2d6","01a06b6b-2720-7258-97c4-699075e3bf22","Seedance 2.0 Fast 提供快速视频生成功能。它可根据文本提示词生成 4 至 15 秒的视频，并支持多种宽高比、音频生成和增强型网页搜索功能。",{"id":1251,"groupId":517,"locale":934,"slug":518,"title":1252,"pageType":17,"contentSource":24,"contentRef":1253,"sort":15,"status":16,"icon":384,"description":1254,"httpMethod":64,"hasToc":27},"01a06b45-bf16-7b49-a58d-bf34c4e8673d","Seedance 2.0 文生视频","01a06b6b-2d4b-7727-a82b-c696192af9e5","Seedance 2.0 Text-to-Video 提供最高的画面质量。它可根据文本提示词生成 4 至 15 秒的视频，并支持多种宽高比、音频生成和增强型网页搜索功能。",{"id":1256,"groupId":524,"locale":934,"slug":525,"title":1257,"pageType":17,"contentSource":24,"contentRef":1258,"sort":15,"status":16,"icon":384,"description":1259,"httpMethod":64,"hasToc":27},"01a06b45-c0cd-728e-8ec9-552eba5ff85b","Seedance 2.0 图生视频","01a06b6b-381c-7513-a2ee-a8b49dd25d41","Seedance 2.0 Image-to-Video 提供最高的画面质量。它可根据文本提示词生成 4 至 15 秒的视频，并支持多种宽高比、音频生成和增强型网页搜索功能。",{"id":1261,"groupId":531,"locale":934,"slug":532,"title":1262,"pageType":17,"contentSource":24,"contentRef":1263,"sort":15,"status":16,"icon":384,"description":1264,"httpMethod":64,"hasToc":27},"01a06b45-c1a7-7ee2-a262-2b714da0c34a","Seedance 2.0 Fast 文生视频","01a06b6b-3da3-795e-a9c7-3991dd0c4e6c","Seedance 2.0 Fast Text-to-Video 提供最高的画面质量。它可根据文本提示词生成 4 至 15 秒的视频，并支持多种宽高比、音频生成和增强型网页搜索功能。",{"id":1266,"groupId":538,"locale":934,"slug":539,"title":1267,"pageType":17,"contentSource":24,"contentRef":1268,"sort":15,"status":16,"icon":384,"description":1269,"httpMethod":64,"hasToc":27},"01a06b45-c2ca-7975-8a57-ce3fd16f790e","Seedance 2.0 Fast 图生视频","01a06b6b-4330-78ae-a3dc-8b98ead9e1a2","Seedance 2.0 Fast Image-to-Video 提供最高的画面质量。它可根据文本提示词生成 4 至 15 秒的视频，并支持多种宽高比、音频生成和增强型网页搜索功能。",{"id":1271,"groupId":545,"locale":934,"slug":546,"title":505,"pageType":17,"contentSource":24,"contentRef":1272,"sort":15,"status":16,"icon":384,"description":1273,"httpMethod":64,"hasToc":27},"01a06b45-cb2b-7df4-ac8f-1c771b74f489","01a06b6b-73ae-7335-a51f-c7e631e22901","Seedance 2.0 提供最高的画面质量。它可根据文本提示词生成 4 至 15 秒的视频，并支持多种宽高比、音频生成和增强型网页搜索功能。",{"id":1275,"groupId":551,"locale":934,"slug":10,"title":552,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":384,"description":1276,"hasToc":55,"children":1277},"01a06b45-36c1-766d-a4fb-fdd1b5f77437","Seedance 2.0 Mini 将字节跳动（ByteDance）前沿的多模态视频生成技术带入对生成速度与成本效益有极高要求的生产场景中。作为 Seedance 2.0 的轻量化高效版本，Mini 模型在大幅提升推理速度、显著降低单条视频生成成本的同时，完整保留了 2.0 系列的核心多模态生成与镜头控制能力。它与标准版 Seedance 2.0 完全共享相同的 API 接口与调用逻辑，无需更改任何代码即可零成本平滑切换。",[1278,1282,1287],{"id":1279,"groupId":557,"locale":934,"slug":558,"title":552,"pageType":17,"contentSource":24,"contentRef":1280,"sort":15,"status":16,"icon":384,"description":1281,"httpMethod":64,"hasToc":27},"01a06b45-bd61-72f8-ba60-021a61ee6ba6","01a06b6b-2178-719c-bcad-5724add1f48c","Seedance 2.0 Mini 提供快速视频生成功能。它可根据文本提示词生成 4 至 15 秒的视频，并支持多种宽高比、音频生成和增强型网页搜索功能。参考图片宽度限制300px-6000px。",{"id":1283,"groupId":563,"locale":934,"slug":564,"title":1284,"pageType":17,"contentSource":24,"contentRef":1285,"sort":15,"status":16,"icon":384,"description":1286,"httpMethod":64,"hasToc":27},"01a06b45-bff2-7c89-a5ad-54eda4eeb85b","Seedance 2.0 Mini 文生视频","01a06b6b-32ba-7695-90bb-3e060a2a0138","Seedance 2.0 Mini Text-to-Video 提供快速视频生成功能。它可根据文本提示词生成 4 至 15 秒的视频，并支持多种宽高比、音频生成和增强型网页搜索功能。参考图片宽度限制300px-6000px。",{"id":1288,"groupId":570,"locale":934,"slug":571,"title":1289,"pageType":17,"contentSource":24,"contentRef":1290,"sort":15,"status":16,"icon":384,"description":1281,"httpMethod":64,"hasToc":27},"01a06b45-c3a7-7002-8b2c-fd87be06217b","Seedance 2.0 Mini 图生视频","01a06b6b-48d3-7818-bf71-acd1e1a5517e",{"id":1292,"groupId":577,"locale":934,"slug":10,"title":1293,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":579,"description":1294,"hasToc":55,"children":1295},"01a06b45-4970-754f-baa2-d698bbb374b2","快乐马","平台现已正式接入阿里巴巴（Alibaba）最新升级的 Happy Horse 1.1 AI 视频生成大模型 API。作为 1.0 榜首模型的重大迭代版本，1.1 版在保持音视频单次联合生成（Single-pass generation）与多语言精准唇形同步的核心优势外，大幅提升了运镜平滑度、复杂的画面连贯性以及近景人像的皮肤真实质感。模型原生支持文生视频、图生视频及全新的多图参考控制（支持最多 9 张参考图组合），可单次生成 3 至 15 秒、最高 1080p 分辨率的商业级视频。",[1296,1300,1304,1309],{"id":1297,"groupId":584,"locale":934,"slug":585,"title":586,"pageType":17,"contentSource":24,"contentRef":1298,"sort":15,"status":16,"icon":579,"description":1299,"httpMethod":64,"hasToc":27},"01a06b45-9bfd-7b5b-bdd6-d2a8853e23b8","01a06b6a-60f1-7eef-ae26-1eda0e15cf64","HappyHorse 1.1 Reference-to-Video 是阿里巴巴推出的新一代 AI 角色与多图参考生视频大模型。模型支持上传 1–9 张参考图（涵盖角色形象、服装道具、产品外观与场景风格），实现高精度的多图融合与主体锁定，有效解决跨镜头画面漂移问题。可生成 720P\u002F1080P 原生高清视频，最长支持 15 秒渲染。",{"id":1301,"groupId":591,"locale":934,"slug":592,"title":593,"pageType":17,"contentSource":24,"contentRef":1302,"sort":15,"status":16,"icon":579,"description":1303,"httpMethod":64,"hasToc":27},"01a06b45-9cdb-7775-bab9-cf199fca5ca4","01a06b6a-660a-7ce7-943a-066b483e39a3","HappyHorse 1.1 Reference-to-Video Spicy（高动态\u002F高表现力版）是阿里巴巴推出的HappyHorse 系列衍生大模型。在继承多图参考主体锁定与音画同步能力的基础上，Spicy模式是无审查模式且针对高强度动作、激进运镜、强视觉冲击与夸张特效进行了深度微调。模型能够基于参考图驱动大幅度、高张力的动作姿态演变与复杂镜头调度，原生支持多语言角色口型对齐与环境音效，广泛适用于高能动作短剧、爆款短视频、游戏 CG 动效与强视觉商业广告制作。",{"id":1305,"groupId":612,"locale":934,"slug":613,"title":1306,"pageType":17,"contentSource":24,"contentRef":1307,"sort":15,"status":16,"icon":579,"description":1308,"httpMethod":64,"hasToc":27},"01a06b45-b9ce-7c33-8f4a-385f7521a750","快乐马 1.1 文生视频","01a06b6b-09cf-785c-a28b-0a0d6623c5cd","HappyHorse 1.1 图生视频 是阿里巴巴推出的新一代 AI 文生视频大模型。模型采用音视频一体化联合生成架构，可由文本直接生成 720P\u002F1080P 原生高清视频，最长支持 15 秒渲染。原生支持多语言角色口型对齐、环境音效与音画同步（无需二次配音），具备出色的动作流畅度、主体一致性与运镜调度能力，广泛适用于短剧创作、商业广告与社交媒体视频制作。",{"id":1310,"groupId":619,"locale":934,"slug":620,"title":1311,"pageType":17,"contentSource":24,"contentRef":1312,"sort":15,"status":16,"icon":579,"description":1313,"httpMethod":64,"hasToc":27},"01a06b45-bacd-71e7-bf32-509f0b3dd229","快乐马 1.1 图生视频","01a06b6b-0fea-7c11-a400-b6b8f0568c1b","HappyHorse 1.1 图生视频 是阿里巴巴推出的新一代 AI 图生视频大模型。支持单图首帧驱动、首尾帧过渡与短视频动态续写，可生成 720P\u002F1080P 原生高清视频，最长支持 15 秒渲染。采用音视频联合生成架构，原生支持多语言角色口型对齐、环境音效与音画同步。具备出色的动作流畅度、主体一致性与运镜调度能力，广泛适用于电商动态展示、短剧特效与社交媒体制作。",{"id":1315,"groupId":626,"locale":934,"slug":10,"title":627,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":155,"description":1316,"hasToc":55,"children":1317},"01a06b45-4403-7330-b5a0-b2e9fdd8023f","平台现已正式接入 Google 旗舰级 AI 视频生成与多轮编辑大模型 Gemini Omni Flash API。作为替代传统视频模型的突破性产品，Omni Flash 具备原生物理规律理解与真实光影建模能力，支持通过文本、单张\u002F多张图片、音频及已有视频素材输入，单次直出包含原生音效的商业级视频。模型突破性地支持“对话式视频编辑”（Video-to-Video Multi-Turn Editing），开发者与创作团队可通过自然语言指令轻松调整镜头视角、替换背景与人物、修改光影风格并保持长场景的高度连贯性。",[1318,1322],{"id":1319,"groupId":632,"locale":934,"slug":633,"title":634,"pageType":17,"contentSource":24,"contentRef":1320,"sort":15,"status":16,"icon":155,"description":1321,"httpMethod":64,"hasToc":27},"01a079c3-3aa3-7cc8-ab72-513b19ee711c","01a06c19-20f8-7abc-87c4-e5062b43451b","Gemini Omni Flash Image-to-Video 是 Google DeepMind 推出的新一代全模态动态视频生成模型。模型基于原生 Omni 全模态架构设计，能精准解析文本提示词与输入图像语义，生成极具电影感的 24 FPS 动态视频。支持 16:9 与 9:16 画幅输出，单次可生成 3–10 秒流畅画面（官方原生支持最高 4K 超采样放大，部分平台端点当前开放 720P 档位）。凭借卓越的主体一致性、物理仿真推演与镜头运镜调度能力，广泛适用于短剧创作、商业广告、影视特效及社交媒体动效生成。",{"id":1323,"groupId":639,"locale":934,"slug":640,"title":641,"pageType":17,"contentSource":24,"contentRef":1324,"sort":15,"status":16,"icon":155,"description":1325,"httpMethod":64,"hasToc":27},"01a079c3-4c94-7c6e-a506-e8301a9422f6","01a06c19-4202-798a-aa23-9201d502dc3a","Gemini Omni Flash Text-to-Video 是 Google DeepMind 发布的 Gemini Omni 系列首款统一多模态大模型。该模型深度融合 Gemini 的逻辑推理能力与 Veo 的视觉生成技术，支持原生\"任意输入到任意输出\"，能够通过纯文本提示词直接生成带音频的 24 FPS 高质量视频。模型最高支持 1080P 及 4K 超采样输出，支持 16:9 与 9:16 画幅及 3–10 秒单次生成，内置 SynthID 数字水印，并支持多轮对话式视频编辑与上下文保持，适用于短剧创作、商业广告、影视特效及多模态 Agent 工作流。",{"id":1327,"groupId":646,"locale":934,"slug":10,"title":647,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":266,"description":1328,"hasToc":55,"children":1329},"01a06b45-3a3a-726a-9de1-12953d978a51","平台现已正式接入 MiniMax（稀宇科技）推出的 MiniMax H3 (Hailuo 3.0) 旗舰级开源全模态生成大模型 API。H3 摒弃了传统任务分立架构，基于 H3-Omni Transformer 与高压缩比 VAE，实现了文本、图像、视频与音频上下文的统一理解与多模态协同生成。模型可单次直出 4 至 15 秒连贯视频，原生同步输出 32kHz 高保真双声道立体声音效（涵盖对白、环境音与音效）。",[1330,1335,1339],{"id":1331,"groupId":652,"locale":934,"slug":653,"title":1332,"pageType":17,"contentSource":24,"contentRef":1333,"sort":15,"status":16,"icon":266,"description":1334,"httpMethod":64,"hasToc":27},"01a06b45-9505-74df-bd6d-dc098846b481","MiniMax H3 图生视频","01a06b6a-3792-7b84-a360-8c3a66c9707b","MiniMax H3 Image-to-Video（海螺 H3 图生视频）是 MiniMax 推出的新一代全模态 AI 视频大模型。模型支持单图首帧驱动与首尾帧平滑过渡，最高可直出 2K 影视级高清画质，单次生成支持 5~15 秒。基于全新 Omni 全模态统一架构，原生支持音视频联合生成（音效、背景音与多语言角色口型同步），具备出色的镜头运镜控制、物理仿真推演与主体一致性保持，广泛适用于电商动效、广告商用与短剧创作。",{"id":1336,"groupId":659,"locale":934,"slug":660,"title":647,"pageType":17,"contentSource":24,"contentRef":1337,"sort":15,"status":16,"icon":266,"description":1338,"httpMethod":64,"hasToc":27},"01a06b45-95df-7ad1-9f26-1158d6729aac","01a06b6a-3d24-76c3-b489-309c3a40c81a","MiniMax H3：根据文本提示词生成视频，同时保持参考图像中的主体一致。支持 2K 分辨率，视频时长为 5–15 秒。",{"id":1340,"groupId":665,"locale":934,"slug":666,"title":1341,"pageType":17,"contentSource":24,"contentRef":1342,"sort":15,"status":16,"icon":266,"description":1343,"httpMethod":64,"hasToc":27},"01a06b45-995b-764b-91c7-bfa9a34ec8ca","MiniMax H3 文生视频","01a06b6a-52ea-755d-aa04-12d248f6a2d3","MiniMax H3 Text-to-Video（海螺 H3 文生视频）是 MiniMax 推出的新一代 AI 视频生成大模型。基于全新 Omni 全模态架构，模型能精准理解复杂文本 Prompt，最高可直接生成 2K 影视级高清视频，单次最长支持 15 秒渲染。原生集成音视频联合生成（音效、背景音与多语言角色口型同步），具备极致的动作流畅度、真实物理仿真推演与镜头运镜调度能力，广泛适用于商业广告、短剧创作与社交媒体视频制作。",{"id":1345,"groupId":672,"locale":934,"slug":10,"title":673,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":674,"description":1346,"hasToc":55,"children":1347},"01a06b45-326a-7a92-b809-d15f8e344b72","Wan 3.0（万相 3.0）是阿里巴巴推出的新一代全能 AI 视频生成大模型。模型涵盖文生视频与图生视频能力，最高支持 30 秒 1080P\u002F4K 影视级高清画质直出。原生集成音视频联合生成架构（支持音效、背景音乐与多语言口型同步），具备极致的复杂动作物理仿真、场景画面连贯性与多镜头调度控制，可极大地降低画面漂移并提升主体一致性，广泛适用于商业广告、影视特效、短剧创作与社交媒体视觉制作。",[1348,1352,1357,1361],{"id":1349,"groupId":679,"locale":934,"slug":680,"title":681,"pageType":17,"contentSource":24,"contentRef":1350,"sort":15,"status":16,"icon":674,"description":1351,"httpMethod":64,"hasToc":27},"01a06b45-a12c-7a0a-8af0-7298e06645c5","01a06b6a-7cad-7c15-8cde-c5e67d0b6d90","Wan 3.0 Image-to-Video（万相 3.0 图生视频）是阿里巴巴通义万相推出的新一代 AI 图生视频大模型。模型支持单图首帧驱动与首尾帧平滑过渡，最高可直出 30 秒 1080P\u002F4K 影视级高清视频。基于 All-in-One 原生全模态架构，原生集成音视频联合生成（音效、背景音与多语言角色口型对齐），具备出色的镜头运镜控制、高逼真物理仿真与主体一致性保持，广泛适用于电商动效、影视特效、短剧创作与广告商用。",{"id":1353,"groupId":714,"locale":934,"slug":715,"title":1354,"pageType":17,"contentSource":24,"contentRef":1355,"sort":15,"status":16,"icon":674,"description":1356,"httpMethod":64,"hasToc":27},"01a06b45-a72d-7b50-ad6d-879bcc374bab","Wan 3.0 文生视频","01a06b6a-9ac2-766a-8d86-a21c9914174f","Wan 3.0 Text-to-Video（万相 3.0 文生视频）是阿里巴巴通义万相推出的新一代 AI 文生视频大模型。模型能深度理解复杂的自然语言 Prompt，最高可直接生成 30 秒 1080P\u002F4K 影视级高清视频。基于原生音视频联合生成架构（原生支持环境音效、背景音乐与多语言角色口型同步），具备极致的动作流畅度、真实物理仿真推演与镜头运镜调度能力，广泛适用于商业广告、短剧创作、影视特效与社交媒体视频制作。",{"id":1358,"groupId":721,"locale":934,"slug":722,"title":723,"pageType":17,"contentSource":24,"contentRef":1359,"sort":15,"status":16,"icon":674,"description":1360,"httpMethod":64,"hasToc":27},"01a08015-1071-75bc-a98b-e68f32a3655b","01a08017-a752-72c7-8e92-1e694e0e7e65","Wan 3.0 Prime Text-to-Video（万相 3.0 Prime 文生视频）是阿里巴巴通义万相系列推出的极速高品质 AI 文生视频大模型。结合了 Prime 架构的极速渲染推理与 Wan 3.0 全模态底座，模型能深度解析复杂自然语言 Prompt，快速生成最高 30 秒 1080P 高清视频。模型原生集成音视频联合生成（原生支持环境音效、背景音与多语言角色口型同步），在大幅缩短视频生成等待时间的同时，具备极致的动作流畅度、真实物理仿真推演与镜头运镜调度能力，广泛适用于商业广告、短剧快速创作、影视特效与社交媒体视频制作。",{"id":1362,"groupId":728,"locale":934,"slug":729,"title":730,"pageType":17,"contentSource":24,"contentRef":1363,"sort":15,"status":16,"icon":674,"description":1364,"httpMethod":64,"hasToc":27},"01a08015-1142-78c6-8b9d-ea27172e2dbe","01a08017-c462-763d-9d32-db56c1528059","Wan 3.0 Prime Image-to-Video（万相 3.0 Prime 图生视频）是阿里巴巴通义万相系列推出的极速高品质 AI 图生视频大模型。结合 Prime 架构的超高速渲染与 Wan 3.0 的全模态底座，模型支持单图首帧驱动与首尾帧平滑过渡，最高可直出 30 秒 1080P 高清视频。模型原生集成音视频联合生成（环境音效与多语言角色口型同步），在大幅缩短视频渲染等待时间的同时，具备高精度的动作物理仿真、镜头运镜控制与主体一致性保持，广泛适用于电商动态展示、影视特效、短剧快速迭代与商业广告创作。",{"id":1366,"groupId":763,"locale":934,"slug":10,"title":764,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":674,"description":1367,"hasToc":55,"children":1368},"01a06b45-4090-760f-8bbe-dc6ceab1f9e3","平台现已正式提供阿里巴巴通义实验室（Alibaba Tongyi Lab）最新旗舰 AI 视频生成大模型 Wan 2.7 (万相 2.7) 的生产级 API 接入。作为面向专业影视制作与高可控创作管线的重磅升级，Wan 2.7 基于 27B MoE 架构，将视频生成与后期编辑深度融合。模型不仅原生支持文生视频、图生视频与高精度音频同步，更突破性地推出了首尾帧插值控制（First & Last Frame Generation） 与自然语言 Video-to-Video 指令编辑。",[1369,1374,1378],{"id":1370,"groupId":769,"locale":934,"slug":770,"title":1371,"pageType":17,"contentSource":24,"contentRef":1372,"sort":15,"status":16,"icon":231,"description":1373,"httpMethod":64,"hasToc":27},"01a06b45-934c-735e-a103-50e4b250a987","Wan 2.7 文生视频","01a06b6a-2c82-79ad-89e0-5a381d841d63","通义万相 Wan 2.7 Text-to-Video 是阿里云推出的新一代文生视频大模型。模型支持中英文长 Prompt 与智能扩写，可由文本直接生成 720P\u002F1080P 原生高清视频，单次生成最高可达 15 秒。内置原生音画协同与音效同步，支持 16:9、9:16 等多画幅自适应，具备强悍的动作连贯性、真实物理仿真与光影渲染能力，适用于影视广告、短视频与动漫创作。",{"id":1375,"groupId":783,"locale":934,"slug":784,"title":785,"pageType":17,"contentSource":24,"contentRef":1376,"sort":15,"status":16,"icon":231,"description":1377,"httpMethod":64,"hasToc":27},"01a06b45-9b11-7d69-972f-cf67feb3da59","01a06b6a-5bd1-7e45-93f3-a1b38df9bf51","Wan 2.7 图生视频辣味模式可将首帧图像转化为具有电影感的短视频，同时保持稳定的时序细节和富有表现力的角色动作。",{"id":1379,"groupId":790,"locale":934,"slug":791,"title":1380,"pageType":17,"contentSource":24,"contentRef":1381,"sort":15,"status":16,"icon":231,"description":1382,"httpMethod":64,"hasToc":27},"01a06b45-b571-73c7-ae3d-cda8eda2f464","Wan 2.7 图生视频","01a06b6a-efc0-7127-a68f-610574ef8188","通义万相 Wan 2.7 Image-to-Video 是阿里云推出的新一代图生视频模型。支持单图首帧生成、首尾帧过渡控制及短视频动态续写，可生成 720P\u002F1080P 高清视频，单次最高可达 15 秒。具备强悍的镜头运镜与真实物理推演能力，原生支持音频驱动角色口型与动作对齐，并自适应 16:9、9:16 等主流画幅，广泛适用于电商动效、特效创作与影视后期。",{"id":1384,"groupId":797,"locale":934,"slug":10,"title":798,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"icon":799,"description":1385,"hasToc":55,"children":1386},"01a06b45-416a-7266-97b5-3fa58ec0e239","平台现已正式接入快手科技（Kuaishou Technology）推出的 可灵 3.0 (Kling 3.0) 旗舰级视听生成大模型 API。可灵 3.0 基于原生统一的多模态（All-in-One）训练架构，将文生视频、图生视频、参考控场与视频编辑深度集成于单一管线。模型突破了短视频生成限制，支持单次直出 3 至 15 秒高清连贯画面，并内置智能多镜头（Multi-Shot）调度，可自动完成正反打对话与推拉摇移镜头切换。",[1387,1391,1395,1400],{"id":1388,"groupId":804,"locale":934,"slug":805,"title":806,"pageType":17,"contentSource":24,"contentRef":1389,"sort":15,"status":16,"icon":799,"description":1390,"httpMethod":64,"hasToc":27},"01a06b45-9429-74e2-9861-1543af203979","01a06b6a-31ca-7a91-b48e-f1f7b3305af6","Kling 3.0 Omni 提供高质量的文生视频生成，具备流畅的动作、电影级视觉效果、精准的提示词遵循能力以及原生音频，可直接分享成片。它提供开箱即用的 REST 推理 API，性能卓越，无冷启动延迟，且价格实惠。",{"id":1392,"groupId":811,"locale":934,"slug":812,"title":813,"pageType":17,"contentSource":24,"contentRef":1393,"sort":15,"status":16,"icon":799,"description":1394,"httpMethod":64,"hasToc":27},"01a06b45-96bf-7ba4-a8bb-74f614577101","01a06b6a-42e6-7782-95c3-cd8dfaa5e5a1","可灵 Kling Video O1（Omni One）是快手推出的行业首个统一多模态视频大模型。模型将文生视频、图生视频与视频编辑整合于单一引擎，涵盖主体参考、局部重绘、风格重塑、首尾帧过渡及视频续写等全套能力。支持至多 7 个主体联动参考与多视角锁定，可完美维持角色与道具一致性，并支持自然语言对话式“P视频”，适用于全流程影视创作与广告制作。",{"id":1396,"groupId":818,"locale":934,"slug":819,"title":1397,"pageType":17,"contentSource":24,"contentRef":1398,"sort":15,"status":16,"icon":799,"description":1399,"httpMethod":64,"hasToc":27},"01a06b45-b64d-7edd-b871-d24091433580","Kling v3.0 图生视频","01a06b6a-f545-795c-b8b0-bac77763a13a","可灵 Kling v3.0 图生视频（Image-to-Video）是快手推出的新一代全能多模态 AI 视频大模型。基于全新的 Omni 架构，模型以静态图或多图主体为锚点，最高支持 15 秒 4K 影视级画质直出。原生集成音画同出与多语言口型驱动，具备“全能参考”与主体锁定能力以防止画面漂移，并支持智能分镜与复杂运镜调度，广泛适用于广告商用、影视特效与短剧内容创作。",{"id":1401,"groupId":825,"locale":934,"slug":826,"title":1402,"pageType":17,"contentSource":24,"contentRef":1403,"sort":15,"status":16,"icon":799,"description":1404,"httpMethod":64,"hasToc":27},"01a06b45-bc85-7326-afa0-fbe29543e16f","Kling v3.0 文生视频","01a06b6b-1beb-7f7b-9bae-6191a11be904","可灵 Kling v3.0 文生视频（Text-to-Video）是快手推出的新一代 AI 视频生成大模型。基于全新 Omni 原生架构，模型支持中英文复杂 Prompt 精准解析，可直接由文本生成高达 4K 影视级画质视频，单次最长支持 15 秒渲染。原生支持音视频联合生成（音效、背景音与多语言角色口型对齐），具备极致的画面真实感、多镜头连贯调度与复杂物理仿真能力，广泛适用于商业广告、影视特效与自媒体创作。",{"id":1406,"groupId":832,"locale":934,"slug":10,"title":1407,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"topTab":834,"children":1408},"01a06b0c-6650-7927-9c77-505fa3c9492c","接入聊天应用",[1409,1413,1417,1421,1425,1429],{"id":1410,"groupId":838,"locale":934,"slug":839,"title":840,"pageType":32,"contentSource":24,"contentRef":1411,"sort":15,"status":16,"icon":842,"description":1412,"hasToc":27},"01a06b0c-6bc3-78bb-8a5f-7bb1cc3f42b0","01a06b0d-bca1-7d1e-8669-25b98431bf66","在 Workbuddy 中接入 iCreat",{"id":1414,"groupId":846,"locale":934,"slug":847,"title":848,"pageType":32,"contentSource":24,"contentRef":1415,"sort":15,"status":16,"icon":850,"description":1416,"hasToc":27},"01a06b0c-6c95-7842-85fc-0b4275eafad4","01a06b0d-c7fd-7fe8-af5a-40b00fde67eb","配置 Codex 接入 iCreat",{"id":1418,"groupId":854,"locale":934,"slug":855,"title":856,"pageType":32,"contentSource":24,"contentRef":1419,"sort":15,"status":16,"icon":858,"description":1420,"hasToc":27},"01a06b0c-6d75-7116-8c52-dd1208ecca72","01a06b0d-d240-7129-a462-24358ea2a23d","通过 iCreat 配置 Claude Code 桌面应用",{"id":1422,"groupId":862,"locale":934,"slug":863,"title":864,"pageType":32,"contentSource":24,"contentRef":1423,"sort":15,"status":16,"icon":866,"description":1424,"hasToc":27},"01a06b0c-6ea6-7f3a-b9d7-b1258d859a7d","01a06b0d-dc90-7792-9b00-a7c0dbd85c23","在 ChatBox 中通过 OpenAI 兼容 API 接入 iCreat",{"id":1426,"groupId":870,"locale":934,"slug":871,"title":872,"pageType":32,"contentSource":24,"contentRef":1427,"sort":15,"status":16,"icon":874,"description":1428,"hasToc":27},"01a06b0c-6fdb-7b21-bf85-d1b82df832cf","01a06b0d-e6a2-73ab-8988-6f78874d1d9e","在 Cherry Studio 中通过 OpenAI 兼容 API 接入 iCreat",{"id":1430,"groupId":878,"locale":934,"slug":879,"title":880,"pageType":32,"contentSource":24,"contentRef":1431,"sort":15,"status":16,"icon":882,"description":1432,"hasToc":27},"01a06b0c-70af-7519-b6ee-21078449d0aa","01a06b0d-f08f-76c5-bc5a-2c3c2014128d","在 AnythingLLM 中通过 OpenAI 兼容 API 接入 iCreat",{"id":1434,"groupId":886,"locale":934,"slug":10,"title":1435,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"topTab":834,"children":1436},"01a06b0c-6776-7999-bafc-596fe12e1d94","接入开发工具",[1437,1441,1445,1449],{"id":1438,"groupId":891,"locale":934,"slug":892,"title":893,"pageType":32,"contentSource":24,"contentRef":1439,"sort":15,"status":16,"icon":858,"description":1440,"hasToc":27},"01a06b0c-7183-7a39-b189-b9b50dc57c4f","01a06b0d-fac4-79d5-aa1a-979dbaa2eadd","在终端通过 Claude Code CLI 接入 iCreat",{"id":1442,"groupId":898,"locale":934,"slug":899,"title":900,"pageType":32,"contentSource":24,"contentRef":1443,"sort":15,"status":16,"icon":902,"description":1444,"hasToc":27},"01a06b0c-72a7-73f6-8c0c-194eb249d06c","01a06b0e-057f-7c4a-9c51-e342e64ab342","在 Cursor 编辑器中接入 iCreat API",{"id":1446,"groupId":906,"locale":934,"slug":907,"title":908,"pageType":32,"contentSource":24,"contentRef":1447,"sort":15,"status":16,"icon":910,"description":1448,"hasToc":27},"01a06b0c-737c-7d7d-8345-f5ddbd9ec591","01a06b0e-10aa-742f-8e92-2ccb37207160","在 OpenCode 中接入 iCreat API",{"id":1450,"groupId":914,"locale":934,"slug":915,"title":916,"pageType":32,"contentSource":24,"contentRef":1451,"sort":15,"status":16,"icon":918,"description":1452,"hasToc":27},"01a06b0c-74a6-7c4e-9391-b4a03cb3b20b","01a06b0e-1bef-7e85-ab45-ed5e7a5eb6d9","使用 CC-Switch 统一管理 CLI 配置",{"id":1454,"groupId":922,"locale":934,"slug":10,"title":1455,"pageType":12,"contentSource":13,"contentRef":14,"sort":15,"status":16,"topTab":924,"children":1456},"01a06b0c-684a-77b1-95aa-3234a7eb1092","帮助",[1457],{"id":1458,"groupId":928,"locale":934,"slug":924,"title":1459,"pageType":32,"contentSource":24,"contentRef":1460,"sort":15,"status":16,"icon":931,"description":1461,"hasToc":27},"01a06b0c-757d-7aad-b00d-9b2eaa19bd5a","常见问题","01a06b0e-277b-7c83-a72d-2b2dfd03b7a1","常见问题解答",{"locale":934,"slug":410,"title":404,"description":1182,"pageType":17,"markdown":1463,"toc":1464,"navNode":1511},"# Seedream 5.0 Lite\n\nSeedream 5.0 Lite 是字节跳动的图片生成模型，支持文生图与参考图生成。可上传最多 5 张参考图片，输出 2K 或 4K 分辨率，支持 PNG 或 JPG 格式。\n\n:::button\n@label 获取 API Key\n@icon key\n@link \u002Fhub\u002Fkeys\n:::\n:::button\n@label 查看完整文档\n@icon fill-LinkSimple\n@link \u002Fhub\u002Fdocs\u002Fzh\u002Fimage\u002Fbytedance\u002Fseedream-5-0\n:::\n:::button\n@label 复制LLM提示词\n@link #agent-prompt\n:::\n\n## Base URL\n\n```\nhttps:\u002F\u002Fapi.icreat.ai\n```\n\n## 认证\n\n所有 API 请求需要通过 API Key 进行认证。您可以在控制台获取 API Key。\n\n```bash\nexport ICREAT_API_KEY=\"your-api-key-here\"\n```\n\n### HTTP 请求头\n\n```python\nimport os\n\nAPI_KEY = os.environ.get(\"ICREAT_API_KEY\")\nheaders = {\n    \"Content-Type\": \"application\u002Fjson\",\n    \"Authorization\": \"Bearer \" + API_KEY,\n}\n```\n\n> **保护好您的 API Key**\n>\n> 切勿在客户端代码或公开仓库中暴露您的 API Key。请使用环境变量或后端代理。\n\n## 代码示例\n\n图片\u002F视频生成采用**三段式异步调用**：先提交任务获取 `task_id`，再轮询任务状态，最后在状态为 `SUCCEEDED` 时获取生成结果。以下示例使用同一 `task_id` 串联三个步骤。\n\n### 1. 提交任务\n\n向提交接口发送生成请求。成功后响应体返回 `task_id`，请保存该 ID 用于后续轮询与取结果。\n\n:::endpoint POST \u002Fv1\u002Ftask\u002Fsubmit\u002Fbytedance\u002Fseedream-5.0 :::\n\n:::code-panel\n@tab 请求\n@card cURL\n\n```bash\ncurl --fail-with-body --connect-timeout 10 --max-time 60 \\\n  -X POST \"https:\u002F\u002Fapi.icreat.ai\u002Fv1\u002Ftask\u002Fsubmit\u002Fbytedance\u002Fseedream-5.0\" \\\n  -H \"Authorization: Bearer ${ICREAT_API_KEY}\" \\\n  -H \"Content-Type: application\u002Fjson\" \\\n  -d '{\n  \"image\": [\n    \"https:\u002F\u002Fupload.icreat.ai\u002Fupload\u002Fxxx.png\"\n  ],\n  \"prompt\": \"老鹰在飞翔\",\n  \"size\": \"2K\",\n  \"watermark\": false,\n  \"outputFormat\": \"png\"\n}'\n```\n\n```javascript\nconst response = await fetch('https:\u002F\u002Fapi.icreat.ai\u002Fv1\u002Ftask\u002Fsubmit\u002Fbytedance\u002Fseedream-5.0', {\n  method: 'POST',\n  headers: {\n    Authorization: `Bearer ${process.env.ICREAT_API_KEY}`,\n    'Content-Type': 'application\u002Fjson',\n  },\n  body: JSON.stringify({\n    image: ['https:\u002F\u002Fupload.icreat.ai\u002Fupload\u002Fxxx.png'],\n    prompt: '老鹰在飞翔',\n    size: '2K',\n    watermark: false,\n    outputFormat: 'png',\n  }),\n});\n\nconst data = await response.json();\nconsole.log(data);\n```\n\n```python\nimport os\nimport requests\n\nresponse = requests.post(\n    \"https:\u002F\u002Fapi.icreat.ai\u002Fv1\u002Ftask\u002Fsubmit\u002Fbytedance\u002Fseedream-5.0\",\n    headers={\n        \"Authorization\": f\"Bearer {os.environ['ICREAT_API_KEY']}\",\n        \"Content-Type\": \"application\u002Fjson\",\n    },\n    json={\n        \"image\": [\"https:\u002F\u002Fupload.icreat.ai\u002Fupload\u002Fxxx.png\"],\n        \"prompt\": \"老鹰在飞翔\",\n        \"size\": \"2K\",\n        \"watermark\": False,\n        \"outputFormat\": \"png\",\n    },\n)\nprint(response.json())\n```\n\n@tab 响应\n@card 200 OK\n\n```json\n{\n  \"task_id\": \"task-xxx\"\n}\n```\n\n:::\n\n### 2. 轮询状态\n\n使用提交步骤返回的 `task_id` 查询任务进度。响应体仅包含 `status` 字段（`SUBMITTED` | `SUCCEEDED` | `FAILED`）；为 `SUCCEEDED` 时可获取结果，为 `FAILED` 时表示任务失败。\n\n:::endpoint POST \u002Fv1\u002Ftask\u002Fquery-status :::\n\n:::code-panel\n@tab 请求\n@card cURL\n\n```bash\ncurl --fail-with-body --connect-timeout 10 --max-time 60 \\\n  -X POST \"https:\u002F\u002Fapi.icreat.ai\u002Fv1\u002Ftask\u002Fquery-status\" \\\n  -H \"Authorization: Bearer ${ICREAT_API_KEY}\" \\\n  -H \"Content-Type: application\u002Fjson\" \\\n  -d '{\"task_id\": \"task-xxx\"}'\n```\n\n```javascript\nconst response = await fetch('https:\u002F\u002Fapi.icreat.ai\u002Fv1\u002Ftask\u002Fquery-status', {\n  method: 'POST',\n  headers: {\n    Authorization: `Bearer ${process.env.ICREAT_API_KEY}`,\n    'Content-Type': 'application\u002Fjson',\n  },\n  body: JSON.stringify({ task_id: 'task-xxx' }),\n});\n\nconst data = await response.json();\nconsole.log(data);\n```\n\n```python\nimport os\nimport requests\n\nresponse = requests.post(\n    \"https:\u002F\u002Fapi.icreat.ai\u002Fv1\u002Ftask\u002Fquery-status\",\n    headers={\n        \"Authorization\": f\"Bearer {os.environ['ICREAT_API_KEY']}\",\n        \"Content-Type\": \"application\u002Fjson\",\n    },\n    json={\"task_id\": \"task-xxx\"},\n)\nprint(response.json())\n```\n\n@tab 响应\n@card 200 OK\n\n```json\n{\n  \"status\": \"SUCCEEDED\"\n}\n```\n\n:::\n\n### 3. 获取结果\n\n任务成功后，使用同一 `task_id` 拉取最终输出。响应体为资源数组，每项包含 `type`、`url` 与 `download_url`；`url` 用于预览或在线访问，`download_url` 用于下载文件。\n\n:::endpoint POST \u002Fv1\u002Ftask\u002Fget-result :::\n\n:::code-panel\n@tab 请求\n@card cURL\n\n```bash\ncurl --fail-with-body --connect-timeout 10 --max-time 60 \\\n  -X POST \"https:\u002F\u002Fapi.icreat.ai\u002Fv1\u002Ftask\u002Fget-result\" \\\n  -H \"Authorization: Bearer ${ICREAT_API_KEY}\" \\\n  -H \"Content-Type: application\u002Fjson\" \\\n  -d '{\"task_id\": \"task-xxx\"}'\n```\n\n```javascript\nconst response = await fetch('https:\u002F\u002Fapi.icreat.ai\u002Fv1\u002Ftask\u002Fget-result', {\n  method: 'POST',\n  headers: {\n    Authorization: `Bearer ${process.env.ICREAT_API_KEY}`,\n    'Content-Type': 'application\u002Fjson',\n  },\n  body: JSON.stringify({ task_id: 'task-xxx' }),\n});\n\nconst data = await response.json();\nconsole.log(data);\n```\n\n```python\nimport os\nimport requests\n\nresponse = requests.post(\n    \"https:\u002F\u002Fapi.icreat.ai\u002Fv1\u002Ftask\u002Fget-result\",\n    headers={\n        \"Authorization\": f\"Bearer {os.environ['ICREAT_API_KEY']}\",\n        \"Content-Type\": \"application\u002Fjson\",\n    },\n    json={\"task_id\": \"task-xxx\"},\n)\nprint(response.json())\n```\n\n@tab 响应\n@card 200 OK\n\n```json\n[\n  {\n    \"type\": \"Image\",\n    \"url\": \"https:\u002F\u002Fcdn.example.com\u002Foutput\u002Fimage.png\",\n    \"download_url\": \"https:\u002F\u002Fcdn.example.com\u002Foutput\u002Fimage.png?download=1\"\n  }\n]\n```\n\n:::\n\n## 输入参数\n\n### 提交任务 — 输入参数\n\n以下参数在提交任务请求体中被接受。\n\n总计: 5 必填: 3 可选: 2\n\n:::field prompt\n@type string\n@required\n\n图片生成提示词，最多 10,000 个字符。\n:::\n\n:::field size\n@type string\n@required\n@default 2K\n\n输出图片分辨率。\n\n@options 2K 4K\n:::\n\n:::field outputFormat\n@type string\n@required\n@default png\n\n输出图片格式。\n\n@options png jpg\n:::\n\n:::field image\n@type array[string]\n\n参考图片 URL 列表，最多 5 张。支持常见图片格式，单张最大 10 MB。留空时为文生图。\n:::\n\n:::field watermark\n@type boolean\n@default true\n\n是否为生成结果添加水印。\n:::\n\n### 轮询状态 — 输入参数\n\n总计: 1 必填: 1 可选: 0\n\n:::field task_id\n@type string\n@required\n\n提交任务接口返回的任务 ID。\n:::\n\n### 获取结果 — 输入参数\n\n总计: 1 必填: 1 可选: 0\n\n:::field task_id\n@type string\n@required\n\n提交任务接口返回的任务 ID。\n:::\n\n## 输出参数\n\n### 提交任务 — 输出参数\n\n总计: 1\n\n:::field task_id\n@type string\n异步任务的唯一标识。使用此 ID 轮询状态并获取结果。\n:::\n\n### 轮询状态 — 输出参数\n\n总计: 1\n\n:::field status\n@type string\n当前任务状态。当状态为 `SUCCEEDED` 时可获取结果。\n\n@options SUBMITTED SUCCEEDED FAILED\n:::\n\n### 获取结果 — 输出参数\n\n响应体为资源对象数组（`array[object]`）。\n\n总计: 3\n\n:::field type\n@type string\n资源类型，图片模型为 `Image`。\n:::\n\n:::field url\n@type string\n生成图片的预览\u002F访问 URL。\n:::\n\n:::field download_url\n@type string\n生成图片的下载 URL（带 attachment 响应头）。\n:::\n\n## LLM友好的提示词 {#agent-prompt}\n\n以下是一段 **LLM 友好的 Markdown 提示词**，可复制到 Cursor、ChatGPT 等 AI 助手中，帮助 AI 理解本模型的 API 接入方式、调用流程与关键参数。点击「复制LLM提示词」按钮或下方代码块均可复制全文。\n\n```markdown\n# bytedance\u002Fseedream-5.0\n\n> Seedream 5.0 Lite 是字节跳动的图片生成模型，支持文生图与参考图生成。\n\n## 概述\n\n通过 iCreat 三段式异步任务 API 提交生成请求，轮询任务状态后在成功时获取图片资源 URL。\n\n## API 信息\n\n- **Base URL**：`https:\u002F\u002Fapi.icreat.ai`\n- **Submit endpoint (POST)**：`\u002Fv1\u002Ftask\u002Fsubmit\u002Fbytedance\u002Fseedream-5.0`\n- **Poll endpoint (POST)**：`\u002Fv1\u002Ftask\u002Fquery-status`\n- **Get result endpoint (POST)**：`\u002Fv1\u002Ftask\u002Fget-result`\n- **Model ID**：`bytedance\u002Fseedream-5.0`\n- **认证**：`Authorization: Bearer ${ICREAT_API_KEY}`\n\n## 调用流程\n\n1. **Submit**：POST submit 路径，Body 见输入要点；响应 `{ \"task_id\": \"...\" }`\n2. **Poll**：POST `\u002Fv1\u002Ftask\u002Fquery-status`，Body `{ \"task_id\": \"...\" }`；响应**仅** `{ \"status\": \"SUBMITTED|SUCCEEDED|FAILED\" }`\n3. **Get result**：`status` 为 `SUCCEEDED` 时 POST `\u002Fv1\u002Ftask\u002Fget-result`；响应为顶层数组，`type` 为 `Image` 时读取 `url` 或 `download_url`\n\n### 输入要点\n\n- 请求体为顶层扁平字段\n- 必填：`prompt`, `size`, `outputFormat`\n- 可选：`image`, `watermark`\n- `prompt`（必填）：图片生成提示词，最多 10,000 个字符。\n- `image`（可选）：参考图片 URL 列表，最多 5 张。支持常见图片格式，单张最大 10 MB。留空时为文生图。\n- `watermark`（可选）：是否为生成结果添加水印。（默认 `true`）\n\n### 输出要点\n\n- 轮询：只看 `status`\n- 结果：顶层 `[{ \"type\": \"Image\", \"url\": \"...\", \"download_url\": \"...\" }]`\n\n## 注意事项\n\n- 三段式须用同一 `task_id` 串联；不可跳过轮询\n- `FAILED` 为终态，需排查请求参数或参考内容\n```",[1465,1469,1471,1475,1477,1480,1483,1486,1488,1491,1494,1497,1499,1502,1505,1508],{"id":1466,"text":1467,"level":1468},"base-url","Base URL",2,{"id":1470,"text":1470,"level":1468},"认证",{"id":1472,"text":1473,"level":1474},"http-请求头","HTTP 请求头",3,{"id":1476,"text":1476,"level":1468},"代码示例",{"id":1478,"text":1479,"level":1474},"1-提交任务","1. 提交任务",{"id":1481,"text":1482,"level":1474},"2-轮询状态","2. 轮询状态",{"id":1484,"text":1485,"level":1474},"3-获取结果","3. 获取结果",{"id":1487,"text":1487,"level":1468},"输入参数",{"id":1489,"text":1490,"level":1474},"提交任务-输入参数","提交任务 — 输入参数",{"id":1492,"text":1493,"level":1474},"轮询状态-输入参数","轮询状态 — 输入参数",{"id":1495,"text":1496,"level":1474},"获取结果-输入参数","获取结果 — 输入参数",{"id":1498,"text":1498,"level":1468},"输出参数",{"id":1500,"text":1501,"level":1474},"提交任务-输出参数","提交任务 — 输出参数",{"id":1503,"text":1504,"level":1474},"轮询状态-输出参数","轮询状态 — 输出参数",{"id":1506,"text":1507,"level":1474},"获取结果-输出参数","获取结果 — 输出参数",{"id":1509,"text":1510,"level":1468},"agent-prompt","LLM友好的提示词",{"id":1180,"groupId":409,"locale":934,"slug":410,"title":404,"pageType":17,"contentSource":24,"contentRef":1181,"sort":15,"status":16,"icon":384,"description":1182,"httpMethod":64,"hasToc":27}]