OpenJev has been released as an open-source project designed to reproduce the interface pattern used by TypeSafe's closed service, Jev, for runtime-defined semantic decisions. Unlike traditional methods that rely on generating text for a software parser to interpret, OpenJev provides a baseline that reads typed option probabilities directly from the model's logits.
PLUS ULTRAProduct LaunchesOpenJev
OpenJev Released to Enable Local Semantic Decision-Making via Model Logits
PLUS ULTRA by Amenoyomi
The project allows for comparing two methods: reading option probabilities directly without decoding, and asking the model to write its option probabilities as JSON text. In a performance comparison using a frozen Qwen3.5-4B model on an RTX 3090, the direct readout method was significantly faster, completing the array 5.21× faster than the generative baseline.
OpenJev is released under the MIT License and supports running on local hardware, such as an RTX 3090, to perform structured decision-making without relying on external APIs.
PLUS ULTRAby Amenoyomi
Many agent decisions—such as routing a request or verifying if evidence supports a claim—are small, structured choices. While standard chat models can handle these, they typically generate text that software must then parse back into logical statements, creating unnecessary overhead. The "Jev pattern" addresses this by treating these decisions as semantic selections among predefined options defined at runtime.
Instead of a decoding loop where the model predicts and writes tokens one by one, OpenJev reads the probabilities (logits) of the allowed options directly from the model's output layer. By normalizing these scores across only the supplied options, the system can determine the model's choice in a single forward pass. This eliminates the need for answer sentences, JSON repair, or the time-consuming process of token generation.
In a system comparison using a frozen Qwen3.5-4B model on an RTX 3090, this direct readout approach proved significantly more efficient. Even when compared to a compact generative baseline that only emitted "yes" or "no" values without keys or explanations, the direct method completed the task 5.21 times faster.
The project is released under the MIT License and is designed for local execution. It requires Python 3.10+, CUDA, and a GPU capable of holding a 4B BF16 model, such as an RTX 3090. To further optimize performance, it supports a shared-state mode that allows a long initial state to be prefetched once and then branched across multiple criteria.
Sources
- OpenJev (Hacker News Frontpage, 2026-09-18)
- GitHub repo ↗