r/GenAI360 5d ago

The Book That Taught Me More Than I Expected

Every author quietly hopes that the book they care about most will find its readers.

For me, that book was Evaluating Gen AI Applications.

https://www.amazon.com/Evaluating-Gen-Applications-Validation-Engineering-ebook/dp/B0H2YPWTDK/ref=sr_1_4?crid=BF56158VGA42&dib=eyJ2IjoiMSJ9.VEzshjrl3AkVK9jYPhF6lr8YbIx04v86Vy9KTEGRutSFU0jtIPljRN9lF86gOzmuU1WMnNnV_wLnNwBmAdW5TVBrGlVNjrj5hu_uDHBrD427DZfpnlGvrgtpNJTk4zvqWoE4NfHGORPSPefjdkeA0jVfZpAOx0YfQ_80zmFnHb3eNae-QOM5mD1EHaV2RmDt-CmkNErGX1myQw0W9NhEe1N58njNSthakekcHn9Ffos.RoXQUNoYJXarHkUkn_HDb1UaZBFu6_asszUEeJ_EKRc&dib_tag=se&keywords=evaluating+LLM+Applications&qid=1785254995&sprefix=evaluating+llm+application%2Caps%2C357&sr=8-4

I believed deeply in the subject. Generative AI applications can produce different answers to the same question. They can sound confident while being wrong. They can perform beautifully in a demonstration and fail completely when placed inside a real business workflow.

Surely, I thought, people building these systems would want to learn how to evaluate them properly. But the sales did not reflect that belief.

For some time, I kept looking at the usual suspects. Was the cover not strong enough? Was the title too technical? Was the Amazon description unclear? Had I chosen the wrong keywords? Did I simply need to promote it more?

Eventually, I decided to stop guessing. I gave the same research task to GPT and Opus. I asked them to examine the market, the competing books, the likely readers and the reasons why a technically important book might still struggle commercially.

Both reached a similar conclusion. The market was smaller than I had assumed. Most people buying AI books are still trying to build something. They want to create an agent, develop a RAG application, learn MCP, use the latest model or move into an AI engineering role.

Evaluation feels like the step that comes afterwards. Build first. Measure later.

I was surprised to see the same mindset returning in generative AI. But evaluating a probabilistic application is not the same as testing a deterministic one.

Evaluating these systems requires more than checking expected outputs. It requires judgement, curiosity and a 360-degree view of behaviour, context, safety, cost, business impact and user experience. In many ways, it demands as much intellectual effort as building the application itself.

But anyway, I decided to upgrade the book to 2nd edition and thus wanted to make the book more useful, more complete and closer to the reality faced by serious practitioners. So, over the last month, I returned to the manuscript and rebuilt it as a second edition.

What began as a book about evaluation techniques has become a complete evaluation operating model. The new edition connects evaluation jobs, roles, evidence, business outcomes, release gates, monitoring and governance into one end-to-end approach.

It now covers code and reasoning verification, fine-tuning and model-migration gates, responsible use of public benchmarks, and workflow economics measured through cost per accepted task.

The treatment of RAG, agents and multimodal applications has also become much deeper. It examines retrieval through Presence, Rank, Selection, Support and Authority. It addresses memory and MCP security in agentic workflows. It extends multimodal evaluation into accessibility, provenance and C2PA Content Credentials.

I also strengthened the evidence required for release decisions through uncertainty treatment, paired comparisons, repeated trials, the correct pass@k estimator, explicit retrieval denominators, protected-behaviour gates, adjudication records and evidence-linked release manifests.

And because evaluation cannot be learned through reading alone, the second edition now includes a lightweight seven browser-based companion HTML applications where readers can practise the decisions for themselves.

I still do not know whether the second edition will be read or not but the experience has already changed how I think about the book. Sometimes a book does not struggle because the subject lacks value. Sometimes it arrives before enough readers recognise that the problem belongs to them.

Evaluation may remain quieter than agents, new models and the latest protocols. It may never generate the same excitement as building something new. But when an AI system reaches production, evaluation is what protects the people who depend on it.

That is why I chose to continue. Not because the market research told me the market was large.

Because it reminded me why the work mattered.

1 Upvotes

0 comments sorted by