Inherit the boundary. Own the gate.
By Keith Townsend
A senior practitioner's method scales through people. Its credibility scales only if the evidence base stays under central control. So this lab builds the thing an advisory practice actually needs: a bench of analysts, one vendor each, all running the same assessment instrument over a corpus the curator vets. On Kamiwaza the pitch is that workrooms supply that boundary architecturally rather than through code you write. The loss condition: an analyst reaches into another analyst's vendor, or scores against evidence nobody vetted, while every boundary in the product reports success.
One deployment, one release, two workrooms, one assessment workload, one pinned model. Everything below is measured on that system, not projected to tenant scale.
Do Inherit the evidence boundary instead of writing one. Isolation is a property of the workroom, not of retrieval code your team maintains: 51 adversarial cross-vendor searches returned 241 hits and not one came from the other workroom, with the agent asking for the absent vendor by name throughout.
Do Own the gate. Build a deterministic corpus validator and run it before scoring rather than as a report afterward. It reads the corpus itself, so the party who contaminated it cannot delete their way out of the finding. This is your work on any platform.
Don’t Do not expect a role to express "can run the instrument, cannot touch the evidence." Both arrive on the same role and the vocabulary is a closed set. Price the process, not a permission.
Don’t Do not trust a status field on this deployment. A 201 has meant no index, a 200 no create, DEPLOYED a dead upstream, and terminal_outcome "indexed" zero vectors. Verify by reading state back.
Kamiwaza sponsored this lab. They supplied a demo deployment and an administrative token, and they did not see the production instrument or choose the vendors. Kamiwaza reviewed the findings before publication for factual errors. Dell and Supermicro appear only as corpus material from their published press releases and neither is scored; Dell is part of the practice's vendor network. Kamiwaza was not scored either. The findings below include the ones they would not have chosen.
Every vendor row I publish in the 4+1 AI Infrastructure model is my judgment applied to evidence I picked. That's the product. It's also the ceiling. The obvious fix is a bench of analysts running the same instrument, one vendor each, so why haven't I built it? Because the moment each analyst controls what evidence their own assessment sees, two rows scored against two different sets of documents stop being comparable. Comparability is the entire value of a matrix.
So the thing worth testing isn't the analysts, and it isn't the instrument. It's whether the platform underneath keeps each person inside the evidence a curator approved.
I built it on Kamiwaza. Two workrooms, one vendor corpus in each, a synthetic eight-layer assessment instrument built only from published 4+1 material, and one Kaizen agent per room. A third vendor got named in the questions and loaded nowhere, as a confabulation control. Nine assessment runs, audited by a deterministic claim resolver rather than by my reading.
Then I attacked it. I asked each workroom to assess the vendor it had no documents for, by name, over and over. Reading the answers isn't enough, because a workroom can refuse to answer while still quietly reaching next door. So I counted what the retrieval layer actually handed back. Across 51 adversarial cross-vendor searches, 241 results came back, and not one of them came from the other workroom. 42 of 42 claims were grounded in the asking room's own corpus. Across 75 assessment cells, zero invented capability claims.
The wall sits under the model, not in it
That distinction matters more than the count. A model that behaves is a model that behaved this time. The boundary here sits at the point where evidence gets fetched, below whichever model happens to be serving, which is why it holds when you swap the model out. It's also the part a buyer cannot establish by reading documentation. You have to go under the interface and count.
What it costs
Here's the trade, in my own framework. The Decision Authority Placement Model (DAPM) asks one question: can you take this somewhere else? In the 4+1 stack, Layer 1B, Context Management and Retrieval within the Data Plane, and Layer 2C, where the governance and the agent lifecycle live, are both Ceded to Kamiwaza. The retrieval path, the workroom binding and the isolation model are theirs, and you can't lift them out and run them on another substrate without rebuilding. My published assessment scored them that way back in July from their documentation alone, and this lab doesn't move the grade. What changed is that I now know what I'm getting for it. Ceding authority is a good trade when the thing you ceded to actually holds.
Layer 1C stays Retained, and that's where this gets interesting. Curation policy is the enterprise's. Which documents count as evidence is your call, not the vendor's. So the question becomes whether the platform gives you a way to say it.
The role I needed doesn't exist
Then the part I didn't expect. I wanted a role that means: this person can run the assessment and cannot touch the evidence. That role isn't there. Owner, editor, viewer, and that's the whole list. A viewer is correctly refused everywhere I pushed, four out of four ways in, including the route that skips the document pipeline and writes straight to the vector store. Read-scoped and write-scoped tokens produced identical outcomes, so token scope isn't an authorization boundary here and shouldn't be treated as one.
But a viewer also can't start a conversation, which means a viewer can't run the instrument either. The person operating the assessment is the same person who can rewrite what it reads.
That isn't a Kamiwaza failure, and I want to be fair about it. It's ordinary access control, and the role model is honest about what it offers. It also isn't something a permission can fix, because the capability you'd need to withhold is the capability you have to grant.
So you catch it instead
You enumerate every document in the room, diff it against the set you approved, and you run that check before anything gets scored. Same code, different place in the pipeline, and only one of those is a control. A report telling you the corpus drifted after you published the row is a receipt, not a gate. Restoring a workroom from a clean corpus took about four minutes, which is what makes the remediation practical rather than theoretical.
What I didn't prove: this is one platform, one week, two vendors, and a synthetic instrument. Nothing here says anything about Kamiwaza's performance, its economics, or how the boundary behaves under real concurrent load with dozens of analysts. And I'm not scoring Kamiwaza in the canon on the strength of a sponsored bench.
The verdict travels in one line. Inherit the boundary. Own the gate. Buy the isolation if the substrate genuinely provides it, and stop rebuilding what a platform already does at the retrieval layer. Then write the corpus check yourself, because the one control nobody is going to hand you is the one that says this evidence is the evidence I approved.
Disclosure: Kamiwaza sponsored this lab. They supplied a demo deployment and an administrative token. They did not see the production instrument, choose the vendors, or hold a veto, and they never asked for one. They reviewed the findings for factual errors before publication and confirmed the work is accurate. Dell and Supermicro appear only as corpus material from their own published press releases. Neither is scored, and Dell is part of this practice's vendor network. Kamiwaza isn't scored either. The ruling is mine.
The video walks through the whole thing, including the part where the number I say out loud is wrong and the slide corrects me: https://youtu.be/E-_zPBFsSdo
The lab, the numbers, the audit tooling and the raw detail: https://labs.layer2c.com/labs/evidence-authority
The numbers
Where each layer belongs
| Layer | Placement |
|---|---|
Layer 1B · Retrieval Context Management & Retrieval The canon places every 1B component here at Ceded, and nothing in this lab moves it. The retrieval path, the workroom binding, and the isolation model are Kamiwaza's, and an enterprise cannot lift them out and run them on another substrate without rebuilding. What the lab adds is the return on that cession, measured rather than assumed. Asked 51 times about a vendor whose corpus sits in the neighbouring workroom, retrieval returned 241 hits and every one came from the asking workroom's own corpus. The boundary is enforced at retrieval rather than left to the model's discretion, so the guarantee does not depend on which model is serving. Ceded authority is a good trade only when the thing you ceded to holds. Here it held. | Ceded |
Layer 1C · Pipelines Data Movement & Pipelines The canon marks 1C a gap and Enterprise Responsibility, and the litmus says a vendor providing nothing leaves the layer Retained by default. That is the placement, and the lab sharpens what it means in practice. Curation policy is the enterprise's, and Kamiwaza's Ceded 2C machinery is what enforces it: a viewer was refused on document upload, collection creation, vector database creation, and the direct vector insert that bypasses the document pipeline, the last at a stricter relation than the others. Read-scoped and write-scoped tokens produced identical outcomes, so token scope is not an authorization boundary here and should not be treated as one. What the enterprise cannot do at this layer is express the separation it most wants, because no role distinguishes reading the corpus from writing it. | Retained |
Layer 2C · Reasoning Agentic Infrastructure — The Reasoning Plane The core of the trade, and the canon already called it: ReBAC enforcement and agent lifecycle governance are proprietary and captive. The lab measures the quality of what is being ceded. The agent is an artifact whose instructions and model binding are fixed at creation and travel with it, so isolation falls out of the same primitive: a principal holding editor in one workroom only, pointed directly at the other workroom's application URL, sees zero agents. The role travels from the workroom membership record into the application session intact and every launch is audited. The limit sits in the same place: the vocabulary is Kamiwaza's, WorkroomRole is a closed enum, and the intent an enterprise most wants to state at this layer is the one it cannot. The constraint surfaces legibly rather than failing silently, which is worth more than it sounds. | Ceded |
Layer 3 (+1) · Applications AI Application Layer — The Value Plane Unremarkable and expected. Any vendor shipping an opinionated application scores this way, because the application's access semantics are its own opinion and cannot be lifted out. The canon grades the layer moderate with the Kaizen agent and App Garden at Ceded and the deployment patterns Delegated, and this lab found nothing that moves it. Worth naming for a buyer: the twenty-three tools the shipped agent carries come from the runtime image, agent-level filters had no effect, and the workroom deploy path exposes no options. On this deployment the surface granted nothing a role did not already hold, because reaching any of it requires a role that can already write. That makes it a reliability cost rather than a governance one. | Ceded |
Method and disclosure
Kamiwaza sponsored this lab. They supplied a demo deployment and an administrative token, and they did not see the production instrument or choose the vendors. They received a pre-publication report carrying every finding, every engineering item and the capability request, and reviewed it for factual errors before publication.
Dell and Supermicro appear only as corpus material, taken from their own published press releases. Neither is assessed here and no grade in this lab says anything about either company. Dell is part of the practice's vendor network, which is disclosed here because the documents are theirs, not because they had any involvement.
The corpus, the fetch tooling with its per-document hashes, the assessment runner, the deterministic claim resolver, its self-test, and the corpus validator are all committed. What stays proprietary is the calibrated assessment methodology: the ratified grading rules, worked reference rows, thresholds, and axis weighting. The instrument used here was built only from published 4+1 material and carries none of it, which is why nothing it produced can be read as an assessment.
The essay on this page was rendered from the lab’s frozen structured record by a model writing under this practice’s voice specification, with a deterministic gate checking every figure against the record before publication. The record is the canonical surface: where prose and record disagree, the record rules.
Download the raw lab detail (Markdown)