Jev for Python engineers
Vercel’s Python AI SDK now supports Jev, a new AI model designed for fast, structured decision-making via multiple-choice evaluations, enabling developers to test classifiers without training data.
Jev is a novel AI model that processes data through multiple-choice questions, returning answers paired with confidence levels. Unlike generative models, it specializes in narrow decision-making tasks, prioritizing speed over broad reasoning. Vercel’s Python team has integrated Jev into the latest AI SDK for Python, allowing developers to experiment with it using a new experimental API. The integration simplifies access to Jev’s capabilities, positioning it as a tool for rapid prototyping rather than a replacement for traditional classifiers.
To use Jev via the Python AI SDK, developers must install the SDK using uv add ai and then call the experimental evaluate() API. This function connects directly to Jev’s decision-making framework, requiring an AI Gateway key for authentication. The process involves setting an environment variable, AI_GATEWAY_API_KEY, and following a provided example to structure questions. The API supports three types of questions: ChoiceQuestion for selecting one answer, ScoreQuestion for rating on a custom scale, and NoulQuestion for estimating probabilities.
The Python AI SDK’s implementation of Jev mirrors its official API, offering a minimalist interface with just one primary function, evaluate(), and a few supporting types. This design choice aims to keep the API straightforward, reducing complexity for developers. The function accepts a model, a state (which can be a string or JSON), and a mapping of questions. Responses are returned in structured JSON, conforming to predefined types and choices, making it easier to integrate with existing systems.
A practical example demonstrates Jev’s use as a classifier in a Python REPL agent. The developer previously spent two days training a custom classifier to distinguish between English and Python code, achieving inconsistent results. Testing Jev on the same task showed improved performance, though it still misclassified partial Python expressions like "what's" + " up" as English. This highlights Jev’s potential for rapid experimentation, even if it requires refinement for specific edge cases.