OFICIAL Hugging Face Blog Gadgets · Jul 15, 2026

Welcome Inkling by Thinking Machines

In brief · 4 sentences
Based on Hugging Face Blog · Jul 15, 2026

Hugging Face introduces Inkling, a 1-trillion-parameter open multimodal model by Thinking Machines that processes text, images, and audio with a 1-million-token context window and speculative decoding layers.

Video

Video available

Key points
·
Main topic: welcome Inkling by Thinking Machines.
·
Category affected: gadgets and hardware.
·
Figures mentioned: 0, 45, 5.
·
The information comes from an official source.
·
The next step is to watch availability, pricing and real-world impact.

The useful question is what changes for users, developers or buyers, and whether the announcement stays industry context or becomes something people can actually use.

Inkling is a new open large language model developed by Thinking Machines and released on Hugging Face, featuring native support for text, image, and audio inputs. The model has approximately 1 trillion parameters and a context window of 1 million tokens, enabling it to process and reason across multiple modalities simultaneously. It is available in two variants: a full BF16 version requiring 2 TB of VRAM and a quantized NVFP4 version requiring 600 GB of VRAM. The model includes speculative multi-token prediction layers to accelerate inference speeds.

The architecture of Inkling incorporates several innovations, including relative attention instead of RoPE for positional encoding, hybrid attention combining global and sliding window mechanisms, and a short 1D convolution layer to enhance local feature processing. It uses a mixture-of-experts design with shared experts, where six routed experts and two shared experts are active during inference. The model employs a hierarchical MLP patchifier for vision and a discretized mel spectrogram approach for audio processing, integrating these modalities without separate encoders.

Inkling offers day-zero support across major inference frameworks, including Hugging Face Transformers, SGLang, vLLM, and llama.cpp, facilitating deployment in both local and server environments. The model can be accessed via Hugging Face Inference Providers for serverless use or through quantized versions for local deployment. For direct inference, users can utilize the any-to-any pipeline in Transformers version 5.14.0 or later, with options to adjust reasoning effort levels for different tasks.

For large-scale deployment, Inkling supports distributed inference using frameworks like SGLang and vLLM, which enable sharding across multiple GPUs and provide OpenAI-compatible APIs. SGLang offers custom model implementations for high-speed serving, while vLLM is optimized for production environments. The model can also be deployed using SLURM for cluster-based parallel processing, with key parameters such as tensor parallel size and context window limits configurable to match hardware constraints.

Original source → Deals on Clipraptor.com →
Extracted signals · detected in the story
Welcome InklingThinking Machines. Hugging FaceWhatInklingOverall CapabilitiesArchitecture Inference Support Transformers SGLangRemote InferenceHugging Face Inference Providers LocalInferenceUnsloth Use Cases Agentic045516