OFICIAL SK Telecom Newsroom

데이터 벽(Data Wall)이 바꾸는 AI 시대 경쟁의 법칙

What happened
Based on SK Telecom Newsroom · Sep 06, 2026

AI firms face a shifting competitive landscape as data scarcity reshapes training strategies, pushing companies to secure proprietary datasets and adapt infrastructure for post-training and retrieval workflows.

데이터 벽(Data Wall)이 바꾸는 AI 시대 경쟁의 법칙
SK Telecom Newsroom — SK Telecom
Key points
·
Epoch AI estimates 300 trillion tokens of high-quality public text data may be depleted between 2026 and 2032 depending on training efficiency.
·
Cloudflare will block AI crawlers from ad-supported pages starting September 15, 2026, restricting access to web content for training.
·
SK Telecom plans to invest in a 5GW AI data center by 2029 to support post-training and inference infrastructure needs.
Key numbers
·
Research by Epoch AI estimates human-written high-quality text data at roughly 300 trillion tokens, with depletion possible between 2026 and 2032 depending on training efficiency.
·
SK Telecom is addressing the Data Wall challenge through three pillars: model development focused on Korean-language data scarcity, responsible utilization of proprietary datasets under strict privacy standards, and investment in AI...

AI development is shifting from relying on vast public datasets to prioritizing high-quality, proprietary data as publicly available text nears exhaustion. Research by Epoch AI estimates human-written high-quality text data at roughly 300 trillion tokens, with depletion possible between 2026 and 2032 depending on training efficiency. This scarcity is compounded by emerging access restrictions, as Cloudflare plans to block AI crawlers from ad-supported pages starting September 15, 2026, tightening control over web content usage for AI training.

The limitations of synthetic data have become evident, with studies like those published in Nature highlighting 'model collapse'—where performance degrades across generations when models are trained on their own outputs. This undermines the assumption that generating data internally can replace human-curated datasets, particularly for rare or exceptional patterns critical to model robustness. Companies are now pursuing real-time licensing agreements with media groups such as Axel Springer and News Corp to secure access to curated content for training and citation in AI responses.

For most enterprises, the immediate challenge is not data scarcity but leveraging internal datasets—transaction logs, customer interactions, maintenance records, and expert annotations—that remain underutilized. Rather than expanding data collection, businesses are advised to define their AI applications first and then identify the specific datasets required, often starting with retrieval-augmented generation (RAG) before scaling to fine-tuning. This approach aligns data strategy with practical deployment needs rather than pursuing volume alone.

SK Telecom is addressing the Data Wall challenge through three pillars: model development focused on Korean-language data scarcity, responsible utilization of proprietary datasets under strict privacy standards, and investment in AI infrastructure including a 5GW AI data center by 2029. The company emphasizes that competitive advantage lies not in data volume but in the ability to responsibly transform internal data into actionable insights while ensuring secure, scalable computing environments for post-training and inference operations.

Original source → Deals on Clipraptor.com →