“답하는 AI에서 일하는 AI로”… A.X K2가 그리는 소버린 AI의 미래 – 김태윤 파운데이션 모델 담당 인터뷰
SK Telecom unveiled A.X K2, a 688-billion-parameter foundation model, with improved reasoning, Korean-language knowledge, and agent capabilities. The model introduces a sparse attention architecture and is accompanied by vision-language and audio models for industry and office use.
The useful question is what changes for users, developers or buyers, and whether the announcement stays industry context or becomes something people can actually use.
SK Telecom introduced its in-house foundation model A.X K2, featuring 688 billion parameters, which enhances mathematical and scientific reasoning, Korean-language knowledge, long-form comprehension, and agent capabilities compared to its predecessor A.X K1. Performance improved by 32.2 percentage points across 14 international benchmarks, with long-form comprehension and agent evaluations showing an 83.9 percentage-point increase. The company also launched vision-language and audio models designed to support manufacturing, defense, biotech, office environments, and daily life applications.
Kim Tae-yoon, head of SKT’s foundation model team, highlighted A.X K2’s shift from a question-answering AI to one that plans and executes tasks autonomously. The model achieved a score of 45.8 on the Apex benchmark, up from 1.0 in A.X K1, and recorded 97.1 on AIME26, placing it among top global open-weight models. For Korean-language tasks, it scored 80.5 on KMMLU-Pro and 91.6 on CLIcK, while agent tool-use ability reached 98.0 on τ²-Bench Telecom. Despite its larger scale, A.X K2 maintains efficient inference by activating only 33 billion parameters during reasoning, achieved through improved data quality, proprietary architecture, and post-training refinement.
A.X K2 incorporates SKT’s proprietary Sparse Gated Attention (SGA) architecture, which enhances efficiency and stability when processing long contexts exceeding 32,000 tokens—equivalent to a report or short novel—and improves token throughput by 67.7% compared to A.X K1 for inputs up to 120,000 tokens. This capability supports extended dialogues, tool calls, and document references required for agentic workflows. The model’s development includes vision-language and audio derivatives, such as A.X K2 VL Light-Preview with a custom vision encoder for interpreting diagrams and charts, and audio models like A.X K2 ALM and A.X K2 Raon-Speech, developed with Krafton, to handle speech recognition and analysis in industrial and office settings.
The performance gains in A.X K2 stem from prioritizing high-quality, domain-specific data, proprietary architectures like SGA, and targeted post-training techniques. These improvements are evident in extreme-difficulty benchmarks such as Apex, where scores rose from 1.0 to 45.8, demonstrating deeper reasoning. SKT is piloting A.X K2 in manufacturing sectors like steel and automotive components, where agents trained on historical defect data can diagnose issues and recommend corrective actions. For defense applications, the model will be quantized for deployment in secure on-premises environments, ensuring sensitive data remains within institutional control.