<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>V2 on Systems &amp; Engineering</title><link>https://valery.tech/ai-engineering/evaluation/v2/</link><description>Recent content in V2 on Systems &amp; Engineering</description><generator>Hugo</generator><language>en-US</language><copyright>Copyright (c) 2014-2023</copyright><atom:link href="https://valery.tech/ai-engineering/evaluation/v2/index.xml" rel="self" type="application/rss+xml"/><item><title>00 Refactor</title><link>https://valery.tech/ai-engineering/evaluation/v2/00-refactor/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://valery.tech/ai-engineering/evaluation/v2/00-refactor/</guid><description>&lt;p&gt;The evaluation mechanics are already compatible with the new delivery mindset; the main changes are at the top of the model: what the product commits to, what a delivery change is, and what &amp;ldquo;release readiness&amp;rdquo; means.&lt;/p&gt;</description></item><item><title>Ai Evaluation</title><link>https://valery.tech/ai-engineering/evaluation/v2/ai-evaluation/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://valery.tech/ai-engineering/evaluation/v2/ai-evaluation/</guid><description>&lt;h1 id="ai-evaluation-as-an-iterative-engineering-practice"&gt;AI Evaluation as an Iterative Engineering Practice&lt;/h1&gt;
&lt;h2 id="1-purpose-and-scope"&gt;1. Purpose and scope&lt;/h2&gt;
&lt;p&gt;AI evaluation is the engineering practice through which a team develops trustworthy, scoped, and decision-relevant understanding of AI-product behaviour.&lt;/p&gt;</description></item><item><title>Ai Evaluation Goals</title><link>https://valery.tech/ai-engineering/evaluation/v2/ai-evaluation-goals/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://valery.tech/ai-engineering/evaluation/v2/ai-evaluation-goals/</guid><description>&lt;p&gt;Yes. A single flat goal list mixes three different questions:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Why do we evaluate?&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;What must the evaluation subsystem do?&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Where do findings go?&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;I&amp;rsquo;d keep one top-level purpose and describe it through three views.&lt;/p&gt;</description></item><item><title>Operationalising Product Behaviour with Automated Evaluators</title><link>https://valery.tech/ai-engineering/evaluation/v2/operationalising-product-behaviour-with-automated-evaluators/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://valery.tech/ai-engineering/evaluation/v2/operationalising-product-behaviour-with-automated-evaluators/</guid><description>&lt;h2 id="1-purpose"&gt;1. Purpose&lt;/h2&gt;
&lt;p&gt;Automated evaluators turn explicit product-behaviour criteria into repeatable judgements that can be applied across many executions.&lt;/p&gt;
&lt;p&gt;They provide the bridge from:&lt;/p&gt;</description></item><item><title>Coverage Design Phase</title><link>https://valery.tech/ai-engineering/evaluation/v2/coverage-design-phase/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://valery.tech/ai-engineering/evaluation/v2/coverage-design-phase/</guid><description>&lt;p&gt;We could reframe the existing user-input work as an upstream coverage-design phase, with user inputs as one artifact, and connect it bidirectionally with failure understanding.&lt;/p&gt;</description></item><item><title>Designing Evaluation Coverage and Cases</title><link>https://valery.tech/ai-engineering/evaluation/v2/10-evaluation-coverage/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://valery.tech/ai-engineering/evaluation/v2/10-evaluation-coverage/</guid><description>&lt;h2 id="1-purpose"&gt;1. Purpose&lt;/h2&gt;
&lt;p&gt;Evaluation coverage design identifies which product behaviours and situations must be represented to answer a defined evaluation question.&lt;/p&gt;</description></item><item><title>Failure Understanding: Developing an Evidence-Linked Model of Recurring AI Product Failures</title><link>https://valery.tech/ai-engineering/evaluation/v2/failure-understanding/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://valery.tech/ai-engineering/evaluation/v2/failure-understanding/</guid><description>&lt;h2 id="1-purpose"&gt;1. Purpose&lt;/h2&gt;
&lt;p&gt;AI products can fail at several points in one execution. A poor final response may follow an earlier misunderstanding, loss of relevant information, incorrect intermediate action, invalid state change, or false report of success. Reviewing only the final response can therefore hide both the first observable problem and the way it affects the rest of the execution.&lt;/p&gt;</description></item><item><title>Failures Putting Into Use</title><link>https://valery.tech/ai-engineering/evaluation/v2/failures-putting-into-use/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://valery.tech/ai-engineering/evaluation/v2/failures-putting-into-use/</guid><description>&lt;h1 id="part-one"&gt;part one&lt;/h1&gt;
&lt;p&gt;Yes. The key distinction is:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A trace, incident, correction, or user complaint is raw product feedback. Failure understanding is the method that turns part of that feedback into structured, decision-relevant evidence about product behaviour.&lt;/strong&gt;&lt;/p&gt;</description></item><item><title>Product Improvement System</title><link>https://valery.tech/ai-engineering/evaluation/v2/product-improvement-system/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://valery.tech/ai-engineering/evaluation/v2/product-improvement-system/</guid><description>&lt;h1 id="ai-product-improvement-system"&gt;AI Product Improvement System&lt;/h1&gt;
&lt;h2 id="1-purpose"&gt;1. Purpose&lt;/h2&gt;
&lt;p&gt;The &lt;strong&gt;AI Product Improvement System&lt;/strong&gt; is the operating model through which an organisation discovers product opportunities, decides which solutions deserve a production commitment, delivers those solutions, operates them, and improves them using evidence from deliberate evaluation and real use.&lt;/p&gt;</description></item></channel></rss>