A colorful "flat lay" arrangement of colorful dahlia blooms.

The Protocol Between Evidence and Answer

How a human author and three AIs answered "Can Line Breeding Stabilize Dahlia Flower Color?" without forcing an answer


Copyright © 2026 by Steve K. Lloyd
All Rights Reserved

The Dahlia Question That Started It All

While Joann Hartwell and I were first talking about my appearance on an upcoming episode of the Dig on Dahlias podcast, she handed me a question I could not answer off the top of my head. Can line breeding stabilize dahlia flower color? She offered it lightly, the way growers ask each other things across a bench, and I had the sense right away that it was heavier than it sounded.


So I made a wager with myself before I started. I would not try only to find out whether my research system could produce an answer, although I hoped that it could. I would also try to find out whether the Dahlia Doctor System could refuse to produce an answer if the evidence did not support one.


That distinction turned out to be the whole story.


The companion article is “Can Line Breeding Stabilize Dahlia Flower Color?” It carries the science. What follows is about how that answer was built and, more to the point, about the moments when it almost came out wrong and something stopped it.


I work with three AI systems and a set of protocols I have been refining for a long time, and the honest reason I use them is not the one people assume. I did not assemble three machines to take a vote. I assembled them to keep any one of us, myself included, from moving too quickly from evidence to conclusion, and from conclusion to polished prose.


Joann asked what sounded like a single question. Before I could answer it, the evidence made me divide it into five. That division is where the work actually happened.



The Dig on Dahlias Podcast

This article was prepared in advance of the author’s interview with The Dig on Dahlias podcast hosts Joann Hartwell and Allison Lingbloom. Once the episode featuring Dahlia Doctor’s Steve Lloyd is available for streaming, readers will find a link to that conversation here.



The logo for the podcast "The Dig on Dahlias."

Before the AI: Evidence Already Has a History


People worry about AI inventing things. That worry is reasonable, but it starts too late in the chain. Long before AI enters the process, the evidence itself has already passed through several hands, and each pair of hands could have changed it.

The title page and table of contents from a 1935 scientific paper on dahlia flower pigmentation genetics.

Figure 1: A photographic reproduction of the title page from a 1935 publication that contributed to the Dahlia Doctor article “Can Line Breeding Stabilize Dahlia Flower Color?”


Consider what the sources for a dahlia color question actually look like. One is a genetics paper from 1935, printed on paper, then scanned and turned into a digital file. Another is a modern experimental study, born digital, clean, and searchable from the start. Others are dissertations that may be indexed poorly or difficult to retrieve in full. A few are in Japanese, German, or Italian, and I rely on translation to work with them. These are not equivalent objects, even when they arrive looking like the same tidy PDF on a screen.


Two distinctions matter here and are easy to miss. Some of the evidence is modern and digitally clean, while some is old, scanned, translated, or difficult to retrieve intact. A searchable PDF may be the publisher’s original digital article, a scan of a printed copy, a machine-readable layer placed over that scan, or some combination of these. Those are not interchangeable.


The path a source travels is longer than it looks. A source may move from printed page to scan, from scan to searchable PDF or machine-read text, and sometimes from its original language into translation. From there it enters a structured record, then an authorized claim, and finally a sentence in the article.


Each transformation creates another place where meaning can change, and only the last few stages involve AI at all. Guard only against the machine and most of the places where meaning slips go unwatched.


Making Each Source Answerable


Early in this project, I needed to know exactly what that 1935 paper reported. Not my memory of it, which was already months old, and not an AI model’s paraphrase of it, which I had no reason to trust on a detail that would carry weight. I needed the paper’s own account, and I needed it in a form I could compare with the paper itself.


That is what my knowledge cards, or KCs, are for. A card is a structured record of a single source, made under a fixed protocol so that every source is described in the same way. It preserves the study closely enough that its limits stay visible: what was studied, under what conditions, what was found, and how far those findings can reasonably be carried.


The KC does not replace the paper. It makes the paper answerable, so that later, when a sentence in the article leans on that source, I can trace the sentence back through the card to the source and ask whether the source really supports it.


Here is where the discipline shows. One of my strongest KCs for this project records a dahlia breeding study that selected for vase life, meaning how long a cut flower lasts, across five generations. The researchers moved the whole population toward the target. It is real dahlia evidence, directly measured, and a clean demonstration that repeated selection can work.


A screenshot showing tht etitle page from a contemporary scientific paper studying dahlia vase life.

Figure 2: The title page from one of the peer-reviewed scientific publications from which the Knowledge Cards originate.

Its usefulness for Joann’s question was real but very narrow because the measured trait was vase life, not color. Because the card kept the study’s design and boundaries visible, the later evidence review could admit it for one narrow purpose: to show what a documented selection response looks like. At the same time, the protocol explicitly prohibited using it as evidence that color would respond in the same way.


The KC makes the source's limits visible. The later protocol decides what those limits authorize in this particular article. The study was relevant and it authorized almost nothing, and the only reason I could keep both facts in view was that the card had kept the study's actual shape in front of me.


A screenshot showing the containts of KC-0217, based on a 1935 scientific paper about the genetuc of dahlia flower pigmentation.

Figure 3: A sample of one of the Knowledge Cards used in the production of this article.


Dividing the Authority


There is an obvious question underneath all of this. If one AI can retrieve sources, read them, draft an answer, and then tell me the answer is sound, why use three?


Because I did not want any single participant, machine or human, holding all of those powers at once. A single conversational model can perform every step and then grade its own work. In an unstructured exchange, I may have no clear way to see where retrieval ended and interpretation or confidence began. So I divided the authority.


One model, Gemini, serves as the reference librarian. It has live read-only access to my full research-card corpus, and its job is to search for and return the cards that bear on a question. It works under a retrieval protocol designed to keep it from deciding what the returned evidence means. It surfaces candidates for later review. It is not authorized to decide the conclusion.


A second model, ChatGPT, formed the authorized claims under a frozen protocol that governed what could be asserted and how strongly. It also produced the first draft, but its prose was not allowed to validate itself.


A third model, Claude, challenged the reasoning and edited the writing. But no editorial improvement became authoritative merely because it read better. It still had to survive an audit against the authorized claims.


And I direct all of it, judge it, rework it, and override it. But my own framing was not protected from the evidence either, and that turned out to matter more than I expected.


A screenshot showing Knowledge Card summaries surfaced by the AI "reference librarian," Gemini.

Figure 4. The Gemini AI “reference librarian” searches the Knowledge Card corpus under a restricted retrieval role, returning candidates for later review rather than deciding what they mean.


I did not build this all at once. It grew because each addition addressed a problem in the previous arrangement. I began with ChatGPT doing most of the reasoning. Gemini came next, separating retrieval from interpretation so that the reasoning model was not relying on its own memory of what it had found. Claude followed, adding independent resistance and editorial judgment. Each addition closed a gap in who was authorized to review whom.


The clearest proof that the system restrains me, and not only the machines, came early. When Claude and I first sketched this article, we had a theory of it. The piece would be about how breeding for color differs from breeding for form. That was our spine.


Then the evidence came back, and the review showed that the spine was slightly wrong. The article could not be organized primarily around form versus color. The deeper question was which kind of stability a breeder actually wanted, and whether crossing plants could deliver that kind.


The difference between form and color still did real work in the finished piece. It simply could no longer be the frame that governed everything else. The evidence redirected the theory, and the system let it. No single voice, mine included, simply got its way.


A smaller version of the same thing happened to the framework itself. At one point, we had six candidate meanings of “stable color.” We pressure-tested one of them and dropped it, leaving five.


We also discovered, partway through, that our own roadmap had a hole in it. We had rules for forming claims and rules for auditing finished prose, but nothing governing the bridge between them: the step where authorized claims are allocated into an article’s structure.


So we added that step while the project was underway. The protocols were working documents, revised under pressure when they proved incomplete, not labels applied afterward to make the work look rigorous.


A screenshot showing output from the AI Claude commenting on ChatGPT

Figure 5. The evidence review showed that the original organizing structure was not the one the sources could support.


The Places Where It Nearly Went Wrong


Everything so far was preparation: research into what the article could say, based on the sources we had assembled. This is where the system did the work I built it to do, and where I learned the most about my own blind spots.


Start with a smaller moment, before the first draft existed. Two historical papers on dahlia color inheritance looked, at a glance, as though they confirmed each other. Two sources pointing the same way are worth more than one, so the temptation was to treat them as mutual support and lean on the pair.


But when the claims were being formed, an audit separated broad agreement from actual replication. It found that the second paper did not reproduce the first paper’s specific model. It offered a narrower, partially independent line of support, and no more.


What had appeared to be two independent demonstrations of the same result turned out to be one detailed result and a second, narrower source of support beside it. The pair could carry a broad statement together, but they could not be counted as two independent demonstrations of the specific model.


The apparent corroboration weakened, and the evidence became more exact at the same time. The system refused a strength that was not really there.

A screenshot of the Dahlia Doctor protocol headed "DDS L-3 Claim Formation Record."

Figure 6. Two papers appeared to confirm the same model, but only one had tested its specific structure. The protocol prevented us from using one source to support the other’s conclusions.


That was the warm-up. The real test came at the end, over the prose.


By that point, we had a draft I was proud of. ChatGPT had helped write it, Claude had polished it, and I had gone through it several times for a careful review. It read well. It was, I thought, close to done.


The final review did not clear the article for publication. Instead, it identified eight passages where the writing had drifted beyond what the evidence supported and placed the article on HOLD.


What matters most is how those problems entered the draft. They were not simply cases of a machine hallucinating and a human catching the error. Several of the troublesome wording choices appeared while we were improving the prose.


Efforts to make writing clearer can also make it overstate the evidence, which is why the final review cannot be left only to whoever improved the prose.


Let me give you three of the eight because they escalate in a way that taught me something.


The first one was mine. In describing a red cultivar whose color fades, I changed the example from mild conditions to hot conditions because "hot" read better. That was it. That was my whole reason. It did not cross my mind that I was touching a claim-restricted passage. 


But the study supporting that passage concerned low-temperature fading, not heat. My small edit, made for readability after I had spent weeks with the evidence, invented a scenario the study never tested.


I built the fence and then clumsily walked through it. The final audit caught the mismatch only because the original boundary was still sitting there in the frozen record, where my edit could not reach it. 


The second correction widened the lens. A sentence claimed that no experiment of a certain kind had been found in the published literature. The record supported only a narrower statement: no such experiment was represented in the evidence assembled for this article.


The problem was not a demonstrated falsehood. The sentence might even have been true, but our assembled evidence could not establish it. Its reach exceeded its support. It enlarged a claim about our own closed evidence set into a claim about the entire scientific literature, without anyone intending it, and that is a much bigger assertion.


Nobody flagged it for weeks because it did not look like an error.


The third was the subtlest, and it is the one that changed how I think about all of this. A closing sentence summarized the article as posing five questions and answering four of them. It was elegant. It was also structurally false.


The evidence spoke to all five meanings of stable color, to varying degrees. What was missing was one central breeding experiment, absent not from the field, necessarily, but from the sources we had gathered. 


"Five questions, four answered" did not overstate a single fact. The whole sentence, gracefully built, simply misdescribed what the article had done.


That was when I understood what the system had really been guarding against. The danger was never mostly bad facts. It was plausible-sounding wrongness that reads well and gets waved through for exactly that reason.


A screenshot of the Dahlia Doctor protocol headed "Review of the current article draft."

Figure 7. The mandatory review built into the protocol placed the article on HOLD, identifying eight passages where the prose exceeded the evidence.


What made those corrections possible is central to the whole process. The revised prose was not sent out for another unrestricted opinion. It was audited against a record the editing stage was not authorized to rewrite.


That frozen record contained the approved claims, their required limitations, and the boundaries of the assembled evidence. Those decisions had been made before the final editing pass. A sentence could be made clearer, warmer, or more graceful, but the editing stage could not loosen an earlier claim simply to accommodate wording we liked. The prose had to remain within the authority already granted to it.


I should be careful about what this proves. Polishing prose can introduce scope drift, and an audit against a fixed standard can catch it. It does not follow that every audit will catch it, or that bolting on a third AI is what saved us.


What saved us was the frozen standard. The revised article was measured against a record the editing stage could not reach back and change.


The declaration that the article was finished was also subject to the protocol. At one point, final approval was declared and then withdrawn because the supporting records required by the process had not yet been completed and locked.


The completion rules still applied, even after the system had announced the work was done. I found that oddly reassuring. I wanted to declare victory right at the moment I least wanted to check my work, and the process would not let me.


Underneath all of these corrections ran a single, harder problem, and it is the one I care about most. The honest answer to Joann’s question was going to be unresolved, and there is a particular way an unresolved answer can go wrong without anyone deciding that it should.


As the evidence accumulated, one bounded statement after another noted that the assembled sources did not address a particular part of the question. Each statement could remain supportable on its own while their accumulated effect began to sound like a soft “no.”


The overall impression could become a verdict the evidence never delivered. Genuinely open is not the same as gently pessimistic.


Keeping that difference visible, and refusing to let the weight of many small absences settle into a conclusion, was the subtlest thing the whole apparatus had to do.

A "flat lay" of colorful dahlia blooms, showing a range of flower colors.

Figure 8. The opening question was Can line breeding stabilize dahlia flower color?


An Answer That Stayed Open


So what did all of this produce?


Not a verdict. It did not conclude that line breeding stabilizes dahlia color, and it did not conclude that it can't. What it produced was a sharper question and several narrower conclusions.


"Stable color" turned out to hide five different things a breeder might want: more seedlings in the target range, a parent that transmits its color more reliably, a color that holds up across seasons and conditions, a color that stays consistent across one plant, and a color that survives being propagated by cuttings.


Those are five separate problems, measured in different ways and answered by different evidence.


The article lays out what is actually known about each. It shows that repeated selection can shift the distribution of a measured trait, and it names the one breeding experiment the assembled sources never included.


Refusing to force a yes or no was not the system failing to answer. It was part of the answer: the evidence supported several bounded conclusions, but not the single verdict the original question invited.


That was the wager, and the process kept it.


What Stayed Mine


The lesson I take is a narrow one, on purpose. I’m not claiming that three AI models are safer than one, that machines arguing with each other arrive at truth, or that a protocol elaborate enough removes the chance of error. It plainly does not. I stepped over a boundary myself, weeks into the work, on evidence I had helped assemble.


Still, the process made authority, evidence boundaries, disagreement, and revision more visible and more answerable than the unstructured back-and-forth I had previously used with conversational AI.


It also produced a documented case in which my own framing was corrected, and the apparent strength of the evidence was corrected, and the polished prose was corrected, and even the announcement that the work was finished was withdrawn and made to wait.


The scientific article was built by me and three AI systems working in defined roles. This companion account I wrote myself, with two of them reading over my shoulder. The systems retrieved, tested, challenged, and helped shape the work, but they did not take responsibility for it.


That part stayed mine, which is exactly as it should be, and exactly what I built the thing to protect.


AI Collaboration Transparency


This article, like its companion “Can Line Breeding Stabilize Dahlia Flower Color?” was created collaboratively by the author, a dahlia grower and educator, and multiple AI language models.

The author directed the structure, tone, scope, and emphasis of the piece; supplied all scientific sources; and retained full editorial control over the final text. The AIs assisted with summarizing complex technical material, suggesting phrasing, and organizing relationships among peer-reviewed sources provided by the author. It did not independently select sources or introduce unsupported claims.

All content was carefully reviewed, edited, and refined by the author to ensure scientific accuracy, clarity, and alignment with the Dahlia Doctor approach to evidence-based horticultural education.


Return to Articles