English

PLUS ULTRA

Prompts are Ephemeral: Building Self-Improving Agent Systems through Measurement and Optimization

PLUS ULTRA by Amenoyomi

In a recent engineering perspective, the focus of building reliable AI agents is shifting from manual prompt engineering to automated optimization based on rigorous measurement. While LLMs can appear reliable in isolation, they often fail in subtle or simple ways when deployed in production environments, such as failing to adhere to structured outputs or brand voices.

To achieve reliability, the author emphasizes that prompts should be treated as ephemeral vectors rather than static instructions. Instead of relying on human-curated prompts, engineers should focus on building "pass^k" (pass power k) test suites that measure an agent's ability to follow specific constraints and handle adversarial scenarios.

By establishing a repeatable measurement through these test suites, developers can implement an optimization flywheel. This involves using an LLM-based optimizer to automatically iterate on prompt structures until they converge on an optimal state that satisfies the defined metrics. This cycle, where agents are used to optimize other agents via feedback loops, is presented as the essential path to scaling reliable, self-correcting agentic systems.

PLUS ULTRAby Amenoyomi

Manual prompt adjustments often fail to ensure long-term reliability because LLM behavior is probabilistic and highly sensitive to context. Even small changes, such as renaming a field in a structured output, may provide a temporary fix that eventually breaks as the model's internal state or the surrounding instructions shift. Because adding any new prompt alters the entire "contextual universe" of the agent, human-led curation cannot reliably predict how these changes will affect existing behaviors or scale across production volumes.

To move beyond this fragility, reliability must be established through "pass^k" (pass power k), where tests are run numerous times to create a repeatable, quantitative measure of success. This approach exposes subtle failures that a few manual checks would miss, turning the perception of reliability into a measurable metric. By establishing this baseline, developers can identify exactly where an agent fails—such as failing to maintain a brand voice or succumbing to adversarial inputs—without relying on intuition.

Once a repeatable measurement exists, the process of prompt refinement can be handed to an LLM-based optimizer. This system analyzes why a prompt performed poorly on the test suite and automatically iterates on the text until it converges on an optimal state. To prevent the optimizer from simply overfitting to the test examples, "holdout tests" are used—a separate set of scenarios the optimizer never sees—to verify that the resulting prompt actually generalizes to new, unseen inputs.

This shift redefines the role of the domain expert from a prompt writer to a curator of evaluation criteria. Since the specific textual content of a prompt is an ephemeral vector that can be generated by a machine, the primary value lies in building "golden datasets" of labeled good and bad responses. The reliability of an agentic system is therefore determined not by the skill of the person writing the prompt, but by the quality of the test suites and the self-improving feedback loop used to optimize them.

Sources

  1. Prompts aren’t Real (Hacker News Frontpage, 2026-09-20)