

Vanchin offers a rich and diverse selection of models, integrating various models for your use. You can learn about the model introductions through the list below, and conveniently integrate the model services into your own business.
Name | Category | Introduction | Model Function | Context Window | Max Input Tokens | Max Output Tokens | Default Model Throttling |
MiniMax-M3 | Image-to-Text Video-to-Text Reasoning | MiniMax-M3 is a cutting-edge multimodal model developed by MiniMax, featuring a self-developed MSA sparse attention architecture that supports ultra-long context windows of up to 1M tokens. The model natively integrates text, image, and video input capabilities, achieving international leadership in professional tasks such as Coding and Agent. | Function Call | 1000K | - | 512K | RPM:30 TPM:200000 |
DeepSeek-V4-Flash-0731 | Reasoning | DeepSeek-V4-Flash-0731 is an efficient hybrid expert (MoE) language model launched by DeepSeek. It achieves extremely low inference costs while maintaining high performance. This model supports an ultra-long context window of up to 1 million (1M) tokens and can handle long documents of the size of a trilogy of three volumes at once. It is pre-trained based on over 32T high-quality and diverse tokens of the corpus and is open-source, providing developers and researchers with lightweight, high-throughput real-time interaction and large-scale deployment solutions. | Function Call | 1024K | - | 384K | RPM:60 TPM:600000 |
MiMo-V2.5 | Reasoning | MiMo-V2.5 is a native omnimodal model developed by Xiaomi, featuring powerful agentic capabilities. It is built upon the MiMo-V2-Flash backbone, with dedicated vision and audio encoders integrated into a unified architecture that supports comprehensive understanding of text, images, video, and audio. The model employs a sparse Mixture of Experts (MoE) architecture, with approximately 310 billion total parameters and about 15 billion activated per inference. The model weights are open-sourced globally under the MIT license. | Function Call | 1000K | - | 128K | RPM:60 TPM:600000 |
MiMo-V2.5-Pro | Reasoning | The Xiaomi MiMo-V2.5 series models include MiMo-V2.5, V2.5-Pro, V2.5-TTS Series, and V2.5-ASR. This represents a comprehensive leap from "usable" to "excellent," and will be open-sourced globally. | Function Call | 1000K | - | 128K | RPM:60 TPM:60 |
Kimi-K2.7-Code | Image-to-Text Video-to-Text Reasoning | Kimi-K2.7-Code is Kimi's newest and most intelligent model, with comprehensive improvements across general Agent capabilities, coding, visual understanding, and more. It has achieved industry-leading results on benchmarks such as Humanity's Last Exam (at the doctoral level), SWE-Bench Pro, and DeepSearchQA. It also supports text, image, and video inputs, both thinking and non-thinking modes, as well as conversation and Agent tasks. | Function Call | 256K | - | 32K | RPM:10 TPM:300000 |
KAT-Coder-Pro-V2.5 | Text Generation Reasoning | KAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make modifications, and complete the entire process in the actual repository. At the same time, it seamlessly integrates multiple experts to fully retain the front-end aesthetic generation capability of V2. | Function Call | 256K | - | 80K | RPM:60 TPM:2200000 |
KAT-Coder-Air-V2.5 | Text Generation Reasoning | KAT-Coder-Air V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make modifications, and complete the entire process in the actual repository. At the same time, it seamlessly integrates multiple experts to fully retain the front-end aesthetic generation capability of V2. | Function Call | 256K | - | 80K | RPM:60 TPM:2200000 |
Keye-VL-2.0-30B-A3B | Image-to-Text Video-to-Text Reasoning | Keye-VL-2.0-30B-A3B is the latest generation 30B-level main base model of the self-developed multimodal large language model Keye family (with 3B activation parameters). This model is the first to introduce the DSA (DeepSeek Sparse Attention) sparse attention mechanism into the multimodal understanding scenario, successfully unlocking the deep perception capability of 256K ultra-long context, and achieving almost lossless reasoning ability in the temporal perception of long videos. At the same time, this is also the first time that the Keye series has built an Agent collaboration mechanism, demonstrating solid system-level collaboration and execution potential in complex scenarios such as Code, Tool, and Search. | Function Call | 256K | 256K | 64K | RPM:3 TPM:600000 |
GLM-5.2 | Reasoning | GLM-5.2 is a fifth-generation large language model developed by Z.ai, a leading artificial intelligence company in China. The model supports a 1M context window and offers more flexible thinking intensity control, making it suitable for complex development, mobile full-stack development, code migration, and scientific research reproduction scenarios. | Function Call | 1024K | 1024K | 128K | RPM:50 TPM:250000 |
Qwen3.5-397B-A17B | Image-to-Text Video-to-Text Reasoning | Qwen3.5-397B-A17B is a high-performance open-source model in the Qwen3.5 series, featuring a causal language model combined with a vision encoder. The model has 397 billion total parameters with 17 billion (17B) activated parameters. As a native vision-language foundation model, it integrates breakthroughs in multimodal learning, architectural efficiency improvements, and large-scale reinforcement learning, aiming to provide developers with exceptional capability and efficiency. | Function Call | 256K | 252K | 64K | RPM:30 TPM:300000 |
DeepSeek-V4-Pro | Reasoning | The DeepSeek-V4 series consists of powerful Mixture of Experts (MoE) language models, including DeepSeek-V4-Pro (1.6T total parameters, 49B activated parameters) and DeepSeek-V4-Flash (284B total parameters, 13B activated parameters). Both models support up to 1 million (1M) token context length and are open-source models pre-trained on over 32T high-quality diverse tokens. | Function Call | 1024K | - | 384K | RPM:10 TPM:300000 |
DeepSeek-V4-Flash | Reasoning | The DeepSeek-V4 series consists of powerful Mixture of Experts (MoE) language models, including DeepSeek-V4-Pro (1.6T total parameters, 49B activated parameters) and DeepSeek-V4-Flash (284B total parameters, 13B activated parameters). Both models support up to 1 million (1M) token context length and are open-source models pre-trained on over 32T high-quality diverse tokens. | Function Call | 1024K | - | 384K | RPM:10 TPM:300000 |
Qwen3-VL-235B-A22B-Thinking | Image-to-Text Video-to-Text | Qwen3-VL-235B-A22B-Thinking is a visual understanding model in the Qwen3 series with significantly enhanced multimodal thinking capabilities. The model has been specifically optimized for STEM and mathematical reasoning, with comprehensive improvements in visual perception and recognition, and major upgrades to OCR capabilities. | Function Call | 128K | 124K | 32K | RPM:500 TPM:1000000 |
Qwen3-Coder-Next | Text Generation | Qwen3-Coder-Next is an open-weight language model released by the Qwen team, specifically designed for coding agents and local development. | - | 256K | - | 64K | RPM:5000 TPM:10000000 |
Qwen3-30B-A3B-Thinking-2507 | Reasoning | Qwen3-30B-A3B-Thinking-2507 is an updated version of Qwen3-30B-A3B Thinking Mode. This is a causal language model based on the Mixture-of-Experts (MoE) architecture, with 30.5 billion total parameters but only 3.3 billion activated during model inference. | Function Call | 128K | 124K | 32K | RPM:500 TPM:1000000 |
Qwen3.6-35B-A3B | Image-to-Text Video-to-Text | Qwen3.6-35B-A3B is the first open-weight variant of the Qwen3.6 series. It is a causal language model with a Vision Encoder, utilizing a Mixture of Experts (MoE) architecture. | Function Call | 256K | 254K | 64K | RPM:30 TPM:300000 |
Qwen3.6-27B | Image-to-Text Video-to-Text | A 27B native vision-language Dense model from the Qwen3.6 series, featuring deep thinking, visual understanding, and text generation capabilities. | Function Call | 256K | 254K | 64K | RPM:30 TPM:300000 |
MiniMax-M2.5 | Reasoning | MiniMax-M2.5 is the latest generation AI model from MiniMax, designed to solve complex real-world tasks, with coding and agent capabilities reaching or exceeding the Opus 4.6 level. | Function Call | 200K | - | 128K | RPM:30 TPM:300000 |
Kimi-K2.6 | Image-to-Text Video-to-Text | Kimi-K2.6 is Kimi's latest and most intelligent model, with comprehensive improvements in general Agent, coding, visual understanding, and other integrated capabilities. | Function Call | 256K | - | 256K | RPM:30 TPM:300000 |
GLM-5.1 | Reasoning | GLM-5.1 is a next-generation flagship model designed for agentic engineering, featuring significantly more powerful coding capabilities than its predecessor. | Function Call | 200K | - | 128K | RPM:50 TPM:250000 |
KAT-Coder-Pro-V2 | Text Generation | A high-performance edition designed for complex enterprise projects and SaaS integration. | - | 200K | - | 80K | RPM:5 TPM:300000 |
Kimi-K2.5 | Reasoning | Kimi K2.5 is the most powerful open-source model to date, built upon Kimi K2 with approximately 15T tokens of continued pre-training on mixed visual and text data. | Function Call | 256K | - | 256K | RPM: TPM: |
GLM-5 | Reasoning | GLM-5 is a next-generation large language model released by Z.ai, primarily targeting complex systems engineering and long-horizon agentic tasks. | Function Call | 200K | - | 128K | RPM:50 TPM:250000 |
MiniMax-M2.1 | Reasoning | As a significant upgrade to the M2 version, it not only retains the advantage of high cost-effectiveness but also notably enhances multilingual programming capabilities and systematically introduces Interleaved Thinking. The model focuses on improving usability across various programming languages and office scenarios, dedicated to helping enterprises and individuals achieve an AI-native way of working and living. | Function Call | 200K | - | 128K | RPM:50 TPM:200000 |
GLM-4.7 | Reasoning | GLM-4.7 is the latest AI model launched by Z.ai, positioned as a new “programming companion” for users. This model demonstrates significant improvements in multilingual agentic programming, terminal task operation, tool usage, and complex reasoning capabilities. | Function Call | 200K | - | 128K | RPM:50 TPM:250000 |
DeepSeek-V3.1-Terminus | Reasoning | DeepSeek-V3.1-Terminus is an updated iteration of DeepSeek-V3.1, designed to preserve the model's core capabilities while addressing issues reported by users through targeted fixes and optimizations. This version retains the same model architecture as DeepSeek-V3 and delivers significant enhancements in specific domains. | Function Call | 128K | 96K | 32K | RPM:500 TPM:1000000 |
MiniMax-M2 | Reasoning | MiniMax-M2 is a lightweight, fast, and highly cost-efficient Mixture-of-Experts (MoE) model with 230B total parameters and 10B activated parameters. While maintaining strong general intelligence, it is deeply optimized for coding and agent tasks. With only 10B activated parameters, it delivers end-to-end tool usage performance that developers expect, while its compact size enables easier deployment and scalability. | Function Call | 200K | 72K | 128K | RPM:5 TPM:300000 |
Qwen3-Coder-480B-A35B-Instruct-FP8 | Text Generation | Qwen3-Coder-480B-A35B-Instruct-FP8 is Alibaba's latest open-source code model, featuring 480B total parameters and 35B activated parameters under a Mixture-of-Experts (MoE) architecture. The model natively supports a 256K-token context length and matches Claude Sonnet 4—the current state-of-the-art—in code understanding, generation, and agent capabilities. | Function Call | 256K | 200K | 64K | RPM:500 TPM:100000 |
Qwen3-VL-235B-A22B-Instruct | Image-to-Text Video-to-Text | The Qwen3 series visual understanding model features comprehensive upgrades in visual coding and spatial perception, delivering significantly enhanced visual perception and recognition, support for ultra-long video understanding, and major OCR improvements. | Function Call | 128K | 126K | 32K | RPM:500 TPM:100000 |
DeepSeek-V3.2-Speciale [Retired] | Reasoning | DeepSeek-V3.2-Speciale is an AI model that balances high computational efficiency with exceptional reasoning and agent capabilities. It is built upon three key technical breakthroughs:DeepSeek Sparse Attention (DSA), a scalable reinforcement learning framework, and a large-scale agent task synthesis pipeline. DeepSeek-V3.2-Speciale is the high-compute variant of this series, designed to further push the boundaries of reasoning performance. | - | 160K | 128K | 64K | RPM:60 TPM:100000 |
DeepSeek-V3.2 | Reasoning | DeepSeek-V3.2, as a transitional step toward next-generation architectures, introduces DeepSeek Sparse Attention—a sparse attention mechanism designed to explore and validate optimizations for training and model inference efficiency in long-context scenarios. | Function Call | 128K | 96K | 64K | RPM:500 TPM:1000000 |
Qwen3-235B-A22B-Instruct-2507 | Text Generation | Qwen3-235b-A22b-Instruct-2507 significantly outperforms QwQ and other non-inference models of comparable size on benchmarks covering mathematics, code generation, and logical reasoning. The model also demonstrates substantial improvements in creative writing, role-playing, multi-turn dialogue, and instruction following, delivering markedly stronger general-purpose capabilities than other models of similar size. | Function Call | 128K | 126K | 32K | RPM:500 TPM:1000000 |
Qwen3-30B-A3B-Instruct-2507 | Text Generation | This model significant improvements in general capabilities—including instruction following, logical reasoning, text understanding, mathematics, science, coding, and tool usage—substantially expanded coverage of long-tail knowledge across multiple languages, and markedly better alignment with user preferences on subjective and open-ended tasks, enabling more helpful responses and higher-quality text generation. | Function Call | 128K | 126K | 32K | RPM:500 TPM:1000000 |
Kimi-K2-Instruct-0905 | Text Generation | Kimi K2 is an advanced Mixture-of-Experts (MoE) language model with 32 billion activated parameters and a total of 1 trillion parameters. Trained using the Muon optimizer, Kimi K2 excels in frontier knowledge, reasoning, and coding tasks, and has been carefully optimized for agent capabilities. | Function Call | 256K | 224K | 32K | RPM:60 TPM:100000 |
DeepSeek-V3 | Text Generation | DeepSeek-V3, open-sourced by DeepSeek, features high-quality pretraining, a scalable Mixture-of-Experts (MoE) architecture, and a complete toolchain for inference and deployment—making it well-suited for broad adoption across enterprises, research institutions, and the open-source community. | Function Call | 64K | 56K | 8K | RPM:500 TPM:1000000 |
DeepSeek-V3.1 | Reasoning | By switching the chat template, DeepSeek-V3.1 now supports both reasoning (thinking) and non-reasoning (non-thinking) modes. Through post-training optimization, the model demonstrates significantly improved performance in tool utilization and agent-like tasks. The model also achieves answer quality on par with DeepSeek-R1-0528, while delivering faster response times. | Function Call | 128K | 96K | 32K | RPM:500 TPM:1000000 |
DeepSeek-V3.2-Exp | Reasoning | As a transitional step toward the next-generation architecture, V3.2-Exp builds upon V3.1-Terminus by introducing DeepSeek Sparse Attention—a sparse attention mechanism designed to explore and validate optimizations in training and inference efficiency for long-context scenarios. | Function Call | 128K | 96K | 64K | RPM:500 TPM:1000000 |
DeepSeek-R1-0528 | Reasoning | DeepSeek-20250528 achieves significant breakthroughs in multi-task generalization, code generation, and mathematical reasoning. It natively supports a 128K-token context window, augmented with an optimized memory mechanism to substantially enhance processing of long documents and multi-turn conversations. | Function Call | 128K | 112K | 16K | RPM:500 TPM:1000000 |
Qwen3-235B-A22B-Thinking-2507 | Reasoning | It significantly outperforms QwQ and non-inference models of comparable size on benchmarks covering mathematics, code, and logical reasoning, achieving state-of-the-art performance among models of its scale. The model achieves industry-leading results in both inference and non-inference modes and supports precise invocation of external tools. | Function Call | 128K | 124K | 32K | RPM:500 TPM:1000000 |
Qwen3-30B-A3B | Text Generation | A compact Mixture-of-Experts (MoE) model with approximately 30 billion total parameters and 3 billion activated parameters. | Function Call | 32K | 30K | 8K | RPM:500 TPM:500000 |
KAT-Coder-Pro V1 | Text Generation | KAT-Coder-Pro V1 possesses advanced intelligent agent capabilities such as multi-tool parallel invocation, enabling autonomous completion of complex tasks with fewer interactions, featuring stronger code comprehension and logical reasoning, delivering ultimate performance for AI Coding. | Function Call | 256K | 256K | 32K | RPM:60 TPM:2000000 |
KAT-Coder-Exp-72B-1010 [Retired] | Text Generation | An RL innovative experimental version within the KAT-Coder series of models. | - | 128K | 128K | 32K | RPM:20 TPM:2000000 |
KAT-Coder-Air V1 [Retired] | Text Generation | A lightweight version within the KAT-Coder series of models. | Function Call | 128K | 128K | 32K | RPM:20 TPM:2000000 |