logologo
Document Center
文档中心
AnnouncementModels Overview

Models Overview


Vanchin offers a rich and diverse selection of models, integrating various models for your use. You can learn about the model introductions through the list below, and conveniently integrate the model services into your own business.

Name

Category

Introduction

Model Function

Context Window

Max Input Tokens

Max Output Tokens

Default Model Throttling

MiniMax-M3

Image-to-Text

Video-to-Text

Reasoning

MiniMax-M3 is a cutting-edge multimodal model developed by MiniMax, featuring a self-developed MSA sparse attention architecture that supports ultra-long context windows of up to 1M tokens. The model natively integrates text, image, and video input capabilities, achieving international leadership in professional tasks such as Coding and Agent.

Function Call

1000K

-

512K

RPM:30

TPM:200000

DeepSeek-V4-Flash-0731

Reasoning

DeepSeek-V4-Flash-0731 is an efficient hybrid expert (MoE) language model launched by DeepSeek. It achieves extremely low inference costs while maintaining high performance. This model supports an ultra-long context window of up to 1 million (1M) tokens and can handle long documents of the size of a trilogy of three volumes at once. It is pre-trained based on over 32T high-quality and diverse tokens of the corpus and is open-source, providing developers and researchers with lightweight, high-throughput real-time interaction and large-scale deployment solutions.

Function Call

1024K

-

384K

RPM:60

TPM:600000

MiMo-V2.5

Reasoning

MiMo-V2.5 is a native omnimodal model developed by Xiaomi, featuring powerful agentic capabilities. It is built upon the MiMo-V2-Flash backbone, with dedicated vision and audio encoders integrated into a unified architecture that supports comprehensive understanding of text, images, video, and audio. The model employs a sparse Mixture of Experts (MoE) architecture, with approximately 310 billion total parameters and about 15 billion activated per inference. The model weights are open-sourced globally under the MIT license.

Function Call

1000K

-

128K

RPM:60

TPM:600000

MiMo-V2.5-Pro

Reasoning

The Xiaomi MiMo-V2.5 series models include MiMo-V2.5, V2.5-Pro, V2.5-TTS Series, and V2.5-ASR. This represents a comprehensive leap from "usable" to "excellent," and will be open-sourced globally.

Function Call

1000K

-

128K

RPM:60

TPM:60

Kimi-K2.7-Code

Image-to-Text

Video-to-Text

Reasoning

Kimi-K2.7-Code is Kimi's newest and most intelligent model, with comprehensive improvements across general Agent capabilities, coding, visual understanding, and more. It has achieved industry-leading results on benchmarks such as Humanity's Last Exam (at the doctoral level), SWE-Bench Pro, and DeepSearchQA. It also supports text, image, and video inputs, both thinking and non-thinking modes, as well as conversation and Agent tasks.

Function Call

256K

-

32K

RPM:10

TPM:300000

KAT-Coder-Pro-V2.5

Text Generation

Reasoning

KAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make modifications, and complete the entire process in the actual repository. At the same time, it seamlessly integrates multiple experts to fully retain the front-end aesthetic generation capability of V2.

Function Call

256K

-

80K

RPM:60

TPM:2200000

KAT-Coder-Air-V2.5

Text Generation

Reasoning

KAT-Coder-Air V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make modifications, and complete the entire process in the actual repository. At the same time, it seamlessly integrates multiple experts to fully retain the front-end aesthetic generation capability of V2.

Function Call

256K

-

80K

RPM:60

TPM:2200000

Keye-VL-2.0-30B-A3B

Image-to-Text

Video-to-Text

Reasoning

Keye-VL-2.0-30B-A3B is the latest generation 30B-level main base model of the self-developed multimodal large language model Keye family (with 3B activation parameters). This model is the first to introduce the DSA (DeepSeek Sparse Attention) sparse attention mechanism into the multimodal understanding scenario, successfully unlocking the deep perception capability of 256K ultra-long context, and achieving almost lossless reasoning ability in the temporal perception of long videos. At the same time, this is also the first time that the Keye series has built an Agent collaboration mechanism, demonstrating solid system-level collaboration and execution potential in complex scenarios such as Code, Tool, and Search.

Function Call

256K

256K

64K

RPM:3

TPM:600000

GLM-5.2

Reasoning

GLM-5.2 is a fifth-generation large language model developed by Z.ai, a leading artificial intelligence company in China. The model supports a 1M context window and offers more flexible thinking intensity control, making it suitable for complex development, mobile full-stack development, code migration, and scientific research reproduction scenarios.

Function Call

1024K

1024K

128K

RPM:50

TPM:250000

Qwen3.5-397B-A17B

Image-to-Text

Video-to-Text

Reasoning

Qwen3.5-397B-A17B is a high-performance open-source model in the Qwen3.5 series, featuring a causal language model combined with a vision encoder. The model has 397 billion total parameters with 17 billion (17B) activated parameters. As a native vision-language foundation model, it integrates breakthroughs in multimodal learning, architectural efficiency improvements, and large-scale reinforcement learning, aiming to provide developers with exceptional capability and efficiency.

Function Call

256K

252K

64K

RPM:30

TPM:300000

DeepSeek-V4-Pro

Reasoning

The DeepSeek-V4 series consists of powerful Mixture of Experts (MoE) language models, including DeepSeek-V4-Pro (1.6T total parameters, 49B activated parameters) and DeepSeek-V4-Flash (284B total parameters, 13B activated parameters). Both models support up to 1 million (1M) token context length and are open-source models pre-trained on over 32T high-quality diverse tokens.

Function Call

1024K

-

384K

RPM:10

TPM:300000

DeepSeek-V4-Flash

Reasoning

The DeepSeek-V4 series consists of powerful Mixture of Experts (MoE) language models, including DeepSeek-V4-Pro (1.6T total parameters, 49B activated parameters) and DeepSeek-V4-Flash (284B total parameters, 13B activated parameters). Both models support up to 1 million (1M) token context length and are open-source models pre-trained on over 32T high-quality diverse tokens.

Function Call

1024K

-

384K

RPM:10

TPM:300000

Qwen3-VL-235B-A22B-Thinking

Image-to-Text

Video-to-Text

Qwen3-VL-235B-A22B-Thinking is a visual understanding model in the Qwen3 series with significantly enhanced multimodal thinking capabilities. The model has been specifically optimized for STEM and mathematical reasoning, with comprehensive improvements in visual perception and recognition, and major upgrades to OCR capabilities.

Function Call

128K

124K

32K

RPM:500

TPM:1000000

Qwen3-Coder-Next

Text Generation

Qwen3-Coder-Next is an open-weight language model released by the Qwen team, specifically designed for coding agents and local development.

-

256K

-

64K

RPM:5000

TPM:10000000

Qwen3-30B-A3B-Thinking-2507

Reasoning

Qwen3-30B-A3B-Thinking-2507 is an updated version of Qwen3-30B-A3B Thinking Mode. This is a causal language model based on the Mixture-of-Experts (MoE) architecture, with 30.5 billion total parameters but only 3.3 billion activated during model inference.

Function Call

128K

124K

32K

RPM:500

TPM:1000000

Qwen3.6-35B-A3B

Image-to-Text

Video-to-Text

Qwen3.6-35B-A3B is the first open-weight variant of the Qwen3.6 series. It is a causal language model with a Vision Encoder, utilizing a Mixture of Experts (MoE) architecture.

Function Call

256K

254K

64K

RPM:30

TPM:300000

Qwen3.6-27B

Image-to-Text

Video-to-Text

A 27B native vision-language Dense model from the Qwen3.6 series, featuring deep thinking, visual understanding, and text generation capabilities.

Function Call

256K

254K

64K

RPM:30

TPM:300000

MiniMax-M2.5

Reasoning

MiniMax-M2.5 is the latest generation AI model from MiniMax, designed to solve complex real-world tasks, with coding and agent capabilities reaching or exceeding the Opus 4.6 level.

Function Call

200K

-

128K

RPM:30

TPM:300000

Kimi-K2.6

Image-to-Text

Video-to-Text

Kimi-K2.6 is Kimi's latest and most intelligent model, with comprehensive improvements in general Agent, coding, visual understanding, and other integrated capabilities.

Function Call

256K

-

256K

RPM:30

TPM:300000

GLM-5.1

Reasoning

GLM-5.1 is a next-generation flagship model designed for agentic engineering, featuring significantly more powerful coding capabilities than its predecessor.

Function Call

200K

-

128K

RPM:50

TPM:250000

KAT-Coder-Pro-V2

Text Generation

A high-performance edition designed for complex enterprise projects and SaaS integration.

-

200K

-

80K

RPM:5

TPM:300000

Kimi-K2.5

Reasoning

Kimi K2.5 is the most powerful open-source model to date, built upon Kimi K2 with approximately 15T tokens of continued pre-training on mixed visual and text data.

Function Call

256K

-

256K

RPM:

TPM:

GLM-5

Reasoning

GLM-5 is a next-generation large language model released by Z.ai, primarily targeting complex systems engineering and long-horizon agentic tasks.

Function Call

200K

-

128K

RPM:50

TPM:250000

MiniMax-M2.1

Reasoning

As a significant upgrade to the M2 version, it not only retains the advantage of high cost-effectiveness but also notably enhances multilingual programming capabilities and systematically introduces Interleaved Thinking. The model focuses on improving usability across various programming languages and office scenarios, dedicated to helping enterprises and individuals achieve an AI-native way of working and living.

Function Call

200K

-

128K

RPM:50

TPM:200000

GLM-4.7

Reasoning

GLM-4.7 is the latest AI model launched by Z.ai, positioned as a new “programming companion” for users. This model demonstrates significant improvements in multilingual agentic programming, terminal task operation, tool usage, and complex reasoning capabilities.

Function Call

200K

-

128K

RPM:50

TPM:250000

DeepSeek-V3.1-Terminus

Reasoning

DeepSeek-V3.1-Terminus is an updated iteration of DeepSeek-V3.1, designed to preserve the model's core capabilities while addressing issues reported by users through targeted fixes and optimizations. This version retains the same model architecture as DeepSeek-V3 and delivers significant enhancements in specific domains.

Function Call

128K

96K

32K

RPM:500

TPM:1000000

MiniMax-M2

Reasoning

MiniMax-M2 is a lightweight, fast, and highly cost-efficient Mixture-of-Experts (MoE) model with 230B total parameters and 10B activated parameters. While maintaining strong general intelligence, it is deeply optimized for coding and agent tasks. With only 10B activated parameters, it delivers end-to-end tool usage performance that developers expect, while its compact size enables easier deployment and scalability.

Function Call

200K

72K

128K

RPM:5

TPM:300000

Qwen3-Coder-480B-A35B-Instruct-FP8

Text Generation

Qwen3-Coder-480B-A35B-Instruct-FP8 is Alibaba's latest open-source code model, featuring 480B total parameters and 35B activated parameters under a Mixture-of-Experts (MoE) architecture. The model natively supports a 256K-token context length and matches Claude Sonnet 4—the current state-of-the-art—in code understanding, generation, and agent capabilities.

Function Call

256K

200K

64K

RPM:500

TPM:100000

Qwen3-VL-235B-A22B-Instruct

Image-to-Text

Video-to-Text

The Qwen3 series visual understanding model features comprehensive upgrades in visual coding and spatial perception, delivering significantly enhanced visual perception and recognition, support for ultra-long video understanding, and major OCR improvements.

Function Call

128K

126K

32K

RPM:500

TPM:100000

DeepSeek-V3.2-Speciale

[Retired]

Reasoning

DeepSeek-V3.2-Speciale is an AI model that balances high computational efficiency with exceptional reasoning and agent capabilities. It is built upon three key technical breakthroughs:DeepSeek Sparse Attention (DSA), a scalable reinforcement learning framework, and a large-scale agent task synthesis pipeline. DeepSeek-V3.2-Speciale is the high-compute variant of this series, designed to further push the boundaries of reasoning performance.

-

160K

128K

64K

RPM:60

TPM:100000

DeepSeek-V3.2

Reasoning

DeepSeek-V3.2, as a transitional step toward next-generation architectures, introduces DeepSeek Sparse Attention—a sparse attention mechanism designed to explore and validate optimizations for training and model inference efficiency in long-context scenarios.

Function Call

128K

96K

64K

RPM:500

TPM:1000000

Qwen3-235B-A22B-Instruct-2507

Text Generation

Qwen3-235b-A22b-Instruct-2507 significantly outperforms QwQ and other non-inference models of comparable size on benchmarks covering mathematics, code generation, and logical reasoning. The model also demonstrates substantial improvements in creative writing, role-playing, multi-turn dialogue, and instruction following, delivering markedly stronger general-purpose capabilities than other models of similar size.

Function Call

128K

126K

32K

RPM:500

TPM:1000000

Qwen3-30B-A3B-Instruct-2507

Text Generation

This model significant improvements in general capabilities—including instruction following, logical reasoning, text understanding, mathematics, science, coding, and tool usage—substantially expanded coverage of long-tail knowledge across multiple languages, and markedly better alignment with user preferences on subjective and open-ended tasks, enabling more helpful responses and higher-quality text generation.

Function Call

128K

126K

32K

RPM:500

TPM:1000000

Kimi-K2-Instruct-0905

Text Generation

Kimi K2 is an advanced Mixture-of-Experts (MoE) language model with 32 billion activated parameters and a total of 1 trillion parameters. Trained using the Muon optimizer, Kimi K2 excels in frontier knowledge, reasoning, and coding tasks, and has been carefully optimized for agent capabilities.

Function Call

256K

224K

32K

RPM:60

TPM:100000

DeepSeek-V3

Text Generation

DeepSeek-V3, open-sourced by DeepSeek, features high-quality pretraining, a scalable Mixture-of-Experts (MoE) architecture, and a complete toolchain for inference and deployment—making it well-suited for broad adoption across enterprises, research institutions, and the open-source community.

Function Call

64K

56K

8K

RPM:500

TPM:1000000

DeepSeek-V3.1

Reasoning

By switching the chat template, DeepSeek-V3.1 now supports both reasoning (thinking) and non-reasoning (non-thinking) modes. Through post-training optimization, the model demonstrates significantly improved performance in tool utilization and agent-like tasks. The model also achieves answer quality on par with DeepSeek-R1-0528, while delivering faster response times.

Function Call

128K

96K

32K

RPM:500

TPM:1000000

DeepSeek-V3.2-Exp

Reasoning

As a transitional step toward the next-generation architecture, V3.2-Exp builds upon V3.1-Terminus by introducing DeepSeek Sparse Attention—a sparse attention mechanism designed to explore and validate optimizations in training and inference efficiency for long-context scenarios.

Function Call

128K

96K

64K

RPM:500

TPM:1000000

DeepSeek-R1-0528

Reasoning

DeepSeek-20250528 achieves significant breakthroughs in multi-task generalization, code generation, and mathematical reasoning. It natively supports a 128K-token context window, augmented with an optimized memory mechanism to substantially enhance processing of long documents and multi-turn conversations.

Function Call

128K

112K

16K

RPM:500

TPM:1000000

Qwen3-235B-A22B-Thinking-2507

Reasoning

It significantly outperforms QwQ and non-inference models of comparable size on benchmarks covering mathematics, code, and logical reasoning, achieving state-of-the-art performance among models of its scale. The model achieves industry-leading results in both inference and non-inference modes and supports precise invocation of external tools.

Function Call

128K

124K

32K

RPM:500

TPM:1000000

Qwen3-30B-A3B

Text Generation

A compact Mixture-of-Experts (MoE) model with approximately 30 billion total parameters and 3 billion activated parameters.

Function Call

32K

30K

8K

RPM:500

TPM:500000

KAT-Coder-Pro V1

Text Generation

KAT-Coder-Pro V1 possesses advanced intelligent agent capabilities such as multi-tool parallel invocation, enabling autonomous completion

of complex tasks with fewer interactions, featuring stronger code

comprehension and logical reasoning, delivering ultimate performance for AI Coding.

Function Call

256K

256K

32K

RPM:60

TPM:2000000

KAT-Coder-Exp-72B-1010

[Retired]

Text Generation

An RL innovative experimental version within the KAT-Coder series of models.

-

128K

128K

32K

RPM:20

TPM:2000000

KAT-Coder-Air V1

[Retired]

Text Generation

A lightweight version within the KAT-Coder series of models.

Function Call

128K

128K

32K

RPM:20

TPM:2000000

上一篇:KwaiKAT Model Development Tool Integration Guide下一篇:Limited-Time Bonus for New Users
该篇文档内容是否对您有帮助?
有帮助没帮助