iFlytek Launches Spark Token Factory for Enterprise AI Model Routing and Governance
The platform is aimed at companies managing multiple AI models and includes an inference engine optimized for private deployments using Huawei Ascend hardware.
iFlytek on July 19 launched Spark Token Factory, an enterprise platform designed to provide a unified access layer for AI models, route requests among models, support governance and optimize token costs.
The company positions the product as a middle layer between enterprise applications and large language models. It is intended for deployments spanning coding assistants, customer service, knowledge retrieval, content generation and document processing, where companies may need to connect to multiple model services.
According to iFlytek, Spark Token Factory combines unified model access, intelligent routing, token-cost optimization and security and compliance governance. The company said the platform is designed to create a closed loop covering access, governance, observability and operations, giving enterprises a centralized entry point for managing model services.
iFlytek said its routing system classifies requests into three levels, L1 through L3, using factors including prompt length, predefined and customized rules, number of dialogue turns and session affinity. The platform then matches requests to model resources based on quality, cost, latency, availability and security level.
The approach is intended to send simpler tasks, such as text classification and information extraction, to more cost-effective models, while assigning long-document analysis and complex reasoning to higher-capability models. iFlytek said the average routing-decision latency can be kept below 100 milliseconds.
For private deployments using domestic hardware, Spark Token Factory also includes an inference engine optimized for Huawei Ascend hardware. iFlytek said it has conducted end-to-end engineering optimization for mainstream open-source large models, covering areas from underlying operators to upper-layer scheduling.
The company said its optimization work includes model compression and precision techniques intended to reduce memory use while maintaining controllable accuracy. It also cited work on fusing core inference operators and reorganizing multi-card communication scheduling to improve computing efficiency.
The announcement addresses a growing operational challenge for enterprises moving from isolated AI pilots to multi-business, multi-scenario and multi-team deployments. As organizations use more models, iFlytek argues they need stronger tools for model management, cost control, system stability and security governance.
iFlytek did not disclose pricing, customer deployments, supported model coverage or independent performance testing.
- 科大讯飞发布星火Token Factory,打造企业级AI模型智能路由与治理新底座qbitai.com / Trade / Published JUL 22, 2026 / Accessed JUL 23, 2026
- 趋境科技华东区域总部落地钱江世纪城,五年内建成万卡级高品质 AI Token 工厂qbitai.com / Trade / Published JUL 22, 2026 / Accessed JUL 23, 2026
- 贝壳财经启动“千帆竞发”计划,将征集百位优质创作者共建内容新生态qbitai.com / Trade / Published JUL 22, 2026 / Accessed JUL 23, 2026