<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>V1 on Systems &amp; Engineering</title><link>https://valery.tech/ai-engineering/evaluation/v1/</link><description>Recent content in V1 on Systems &amp; Engineering</description><generator>Hugo</generator><language>en-US</language><copyright>Copyright (c) 2014-2023</copyright><atom:link href="https://valery.tech/ai-engineering/evaluation/v1/index.xml" rel="self" type="application/rss+xml"/><item><title>01 Starting Sequence</title><link>https://valery.tech/ai-engineering/evaluation/v1/01-starting-sequence/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://valery.tech/ai-engineering/evaluation/v1/01-starting-sequence/</guid><description>&lt;p&gt;Based on the attached files, we have &lt;strong&gt;different levels of startup definition for the three loops&lt;/strong&gt;. The material describes methods and interfaces; it does not establish that the corresponding product, tooling, datasets, or processes have already been implemented.&lt;/p&gt;</description></item><item><title>10 User Inputs</title><link>https://valery.tech/ai-engineering/evaluation/v1/10-user-inputs/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://valery.tech/ai-engineering/evaluation/v1/10-user-inputs/</guid><description>&lt;h1 id="building-a-starting-set-of-user-inputs-for-ai-evaluation"&gt;Building a Starting Set of User Inputs for AI Evaluation&lt;/h1&gt;
&lt;p&gt;This guide addresses building a representative set of &lt;strong&gt;user inputs&lt;/strong&gt;. These inputs form one component of the evaluation dataset and define the portion of the query space that should cover the important ways users may interact with the application.&lt;/p&gt;</description></item><item><title>11 Building Balanced Set</title><link>https://valery.tech/ai-engineering/evaluation/v1/11-building-balanced-set/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://valery.tech/ai-engineering/evaluation/v1/11-building-balanced-set/</guid><description>&lt;h2 id="building-a-balanced-starting-evaluation-set"&gt;Building a balanced starting evaluation set&lt;/h2&gt;
&lt;p&gt;A balanced evaluation set represents the parts of the product that matter, in proportions appropriate to their importance and risk. Here we&amp;rsquo;re building &lt;a href="https://valery.tech/ai-engineering/evaluation/v1/10-user-inputs/"&gt;10 User Inputs&lt;/a&gt; user inputs for evaluation set, but also the methods and principles could be applied to other or more broad concepts.&lt;/p&gt;</description></item><item><title>20 Error Analysis</title><link>https://valery.tech/ai-engineering/evaluation/v1/20-error-analysis/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://valery.tech/ai-engineering/evaluation/v1/20-error-analysis/</guid><description>&lt;h1 id="failure-understanding-discovering-and-structuring-how-llm-systems-fail"&gt;Failure Understanding: Discovering and Structuring How LLM Systems Fail&lt;/h1&gt;
&lt;h2 id="intro"&gt;Intro&lt;/h2&gt;
&lt;p&gt;LLM applications rarely fail in only one visible place. A poor final response may originate in an earlier misunderstanding, an incorrect tool call, missing context, stale environment state, or a failure to preserve a user constraint. Evaluating only the final response can therefore conceal both the first visible problem and the way it propagates through the execution.&lt;/p&gt;</description></item><item><title>AI Evaluation as an Engineering Learning System</title><link>https://valery.tech/ai-engineering/evaluation/v1/ai-evaluation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://valery.tech/ai-engineering/evaluation/v1/ai-evaluation/</guid><description>&lt;p&gt;AI evaluation connects product intent with observed system behaviour. It turns human and automated judgment into reusable knowledge and uses that knowledge to produce evidence for product decisions.&lt;/p&gt;</description></item><item><title>AI Evaluation as an Iterative Engineering Practice</title><link>https://valery.tech/ai-engineering/evaluation/v1/ai-evaluation-revised/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://valery.tech/ai-engineering/evaluation/v1/ai-evaluation-revised/</guid><description>&lt;p&gt;AI evaluation connects product intent, observed system behaviour, and a defined decision or knowledge need. It uses deliberate probes and operational observations to produce evidence, develops explicit Quality Understanding, and applies that understanding to product decisions.&lt;/p&gt;</description></item><item><title>Conceptualization</title><link>https://valery.tech/ai-engineering/evaluation/v1/conceptualization/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://valery.tech/ai-engineering/evaluation/v1/conceptualization/</guid><description>&lt;p&gt;I want to build a conceptual movel&amp;ndash;or environment&amp;ndash;within which our evaluations (evals) subsystem will be designed and will operate. Currently, we use evals for specific tasks, and these tasks exist within a broader context. We can use this framework to guide and validate the design of our evals.&lt;/p&gt;</description></item><item><title>Conceptualization 2</title><link>https://valery.tech/ai-engineering/evaluation/v1/conceptualization-2/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://valery.tech/ai-engineering/evaluation/v1/conceptualization-2/</guid><description>&lt;h1 id="evaluation-conceptualization"&gt;Evaluation Conceptualization&lt;/h1&gt;
&lt;h2 id="1-purpose"&gt;1. Purpose&lt;/h2&gt;
&lt;p&gt;An evaluation system exists because the behavior of an AI product cannot be inferred reliably from its specification or implementation alone.&lt;/p&gt;</description></item><item><title>Guide</title><link>https://valery.tech/ai-engineering/evaluation/v1/guide/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://valery.tech/ai-engineering/evaluation/v1/guide/</guid><description>&lt;p&gt;Where to start?&lt;/p&gt;
&lt;p&gt;It could depend on different factors, but first - get to know most base / ambiguous concepts at &lt;a href="https://valery.tech/ai-engineering/evaluation/v1/taxonomy/"&gt;Taxonomy&lt;/a&gt;. For example, what is trace, span, &amp;hellip;. It gives us alignment these required concepts and tools to investigate the domain further.&lt;/p&gt;</description></item><item><title>Evaluation of LLM Workflows and Coding Agents</title><link>https://valery.tech/ai-engineering/evaluation/v1/harness-and-platform/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://valery.tech/ai-engineering/evaluation/v1/harness-and-platform/</guid><description>&lt;p&gt;This page covers evaluation for LLM-based workflows and coding agents.&lt;/p&gt;
&lt;p&gt;The goal is to make agent behavior &lt;strong&gt;observable, repeatable, and comparable&lt;/strong&gt; across prompts, tools, models, orchestration strategies, and workflow designs.&lt;/p&gt;</description></item></channel></rss>