Baseten on Hugging Face Inference Providers 🔥
Baseten has joined Hugging Face’s Inference Provider ecosystem, enabling serverless AI model deployment directly on the Hub with SDK integration for Python and JavaScript.
Baseten, an AI infrastructure platform offering serverless AI and training, is now an official Inference Provider on Hugging Face Hub. This integration allows developers to deploy and use models like Kimi K3 and DeepSeek V4 Flash directly from model pages, with support initially for conversational and text-generation tasks. The collaboration expands the Hub’s serverless inference capabilities, providing broader access to frontier models without extensive setup.
The integration includes support in Hugging Face’s Python and JavaScript SDKs, enabling developers to route requests to Baseten seamlessly. Authentication via Hugging Face tokens automatically directs queries to Baseten, with billing handled either through Baseten’s account or standard provider rates depending on the request type. This setup ensures transparent pricing without additional markup from Hugging Face.
Baseten’s models are also accessible through Hugging Face Agent Harnesses such as Pi, OpenCode, and Hermes Agents, allowing users to integrate Baseten-hosted models into workflows without additional configuration. The integration covers popular open-weight LLMs, with additional task support planned for future updates.
Hugging Face offers free inference with a small quota for signed-in free users, while PRO users receive $2 in monthly inference credits usable across providers. Billing for routed requests aligns with provider API rates, and future revenue-sharing agreements with providers may be introduced.