Resources

The why, in writing.

Four documents we mean every word of — for caretakers, for patients, for anyone deciding whether to trust software with something that matters. Free to read, free to share.

For caretakers

The Caretaker's Field Guide

A between-visits companion for the person holding the rope.

Start here: the part nobody says out loud

You did not sign up for this.

You signed up for a marriage, or a friendship, or a kid, or a parent. You signed up for a person. Depression arrived later, uninvited, and quietly rewrote your job description. Now you are part scheduler, part short-order cook, part night watchman, part translator of silences — and nobody handed you a manual, a training, or so much as a laminated card.

So here is the truth, up front, where it belongs: this is hard. Not "hard but rewarding" in the brochure way. Hard in the way that makes you forget what month it is. Hard in the way that has you crying in the car in a parking lot for no reason you could name.

You are allowed to be tired. Tired is not a moral failure. Tired is the receipt for work you actually did.

This guide will not fix your person. Nothing in your pocket will. What it can do is keep you on your feet — because you are not the treatment, but you are part of the ground the treatment stands on, and ground needs maintenance too.


The flat days

Most of depression is not dramatic. Most of it is flat — a person on a couch, curtains half-drawn, a day that goes nowhere. The flatness is often harder on you than a crisis would be, because a crisis has a script and flatness has nothing. Just hours.

What helps on a flat day:

  • Lower the bar until it's on the floor, then stand on it together. A shower is a win. Toast is a win. Sitting in a different room than the one they woke up in is a win. You are not being condescending by counting these; you are being accurate about what today's terrain actually looks like.
  • Parallel presence beats conversation. You do not have to fill the silence. Fold laundry in the same room. Watch the dumb show. Depression makes talking expensive; being near is a currency they can still afford.
  • Do not audit the day at 6pm. "What did you do today?" lands like an invoice. If nothing happened today, nothing happened today. The day is over either way.

And what helps you: remember that flat is not the same as failing. A flat day where nothing got worse is a day the two of you held the line. Nobody claps for holding the line. Hold it anyway.


Translating "I'm fine"

You will hear "I'm fine" a thousand times, and you will learn — probably already have — that it comes in dialects. There's the "fine" that means fine-ish. The "fine" that means please don't make me produce feelings right now. The "fine" that means I am nowhere near fine and I'm ashamed of it.

You cannot always tell which one you got. You don't have to.

Instead of prosecuting the "fine" ("You don't seem fine—"), try leaving a door open without shoving them through it:

  • "Okay. If it stops being fine, I want to know. Even at a bad hour."
  • "You don't have to talk. Want company anyway?"
  • Or just: "Okay," followed by staying in the room a little longer than you needed to.

The goal is not to extract the truth on demand. The goal is to be the person the truth can find when it's ready to come out. That's a slower job, and a better one.

One more thing: sometimes "I'm fine" is aimed at protecting you, because they can see you're running on fumes. That's worth knowing. It doesn't mean you should hide your tiredness better. It means you're both carrying more than you're saying, and one of you can start telling the truth first. It's allowed to be you.


When they refuse

They will refuse things. The walk. The shower. The appointment. The dinner you made — the one they said yesterday they wanted.

Refusal feels personal. It mostly isn't. Depression sits between your offer and their reach, and from where they are, "no" is often the only word within grabbing distance.

Some ground rules that keep refusals from turning into wars:

  • Offer twice, then let it be. The first offer is an invitation. The second is a door left open. A third becomes pressure, and pressure gives depression something to push back against.
  • Shrink the ask. Not "come for a walk" — "come stand on the porch with me for two minutes." Not "eat dinner" — "eat three bites while I sit with you." You can almost always cut an ask in half. Some days, in quarters.
  • Do not negotiate the non-negotiables alone. Medication, appointments, safety — if these are being refused for real, that is not a couch conversation you're supposed to win by yourself. Loop in their clinician. That's not tattling. That's the whole reason the clinician exists.

And forgive yourself for the flash of anger when the tenth reasonable offer gets the tenth flat "no." You made a good offer. The "no" belongs to the illness more than to you.


Small wins that don't feel like wins

Here is a strange grief nobody warns you about: the wins are so small they can insult you.

They answered a text from a friend. They ate breakfast without being asked. They made a joke — one joke, half a joke — and for a second the old face surfaced. And some part of you thinks: this is what we're celebrating now?

Yes. This is what you're celebrating now. Not because you've lowered your standards, but because you've learned to read the actual scoreboard. Recovery from a bad stretch is not a montage. It's a jagged line that trends upward so slowly you can only see it by looking back weeks, not days. The single joke is data. The unprompted breakfast is movement.

Say the win out loud, lightly, without confetti: "Hey, you texted Sam back. Good." Then move on. Depression hates a spotlight, but it can tolerate a nod.

And write the wins down somewhere, even one line in your phone. Not for them — for you, for the night three weeks from now when you're convinced nothing has changed, and your own notes quietly prove you wrong.


2am

There is a version of this life that only happens after midnight, and if you know it, you know it: the listening through the wall. The checking. The lying still doing arithmetic about how much worry is the right amount of worry. The particular loneliness of being the only person awake in a house that is not okay.

For the 2am hours:

  • Decide what 2am is for, in daylight. At 2am your brain is a terrible executive. So make the decisions at 2pm: what counts as "call someone," who you'd call, what can genuinely wait until morning. Then, at 2am, you're not deciding — you're just following instructions from a saner version of yourself.
  • Most 2am fears are 2pm problems wearing a mask. Not all. But most. If it can be written on a sticky note and dealt with tomorrow, write it down. The note holds the worry so you don't have to hold it all night.
  • If the fear is about their safety — truly about their safety — that is the one fear you never sit alone with. Have the number ready before you need it. This is the single piece of boilerplate this guide will allow itself, because it's true: you were never meant to do the dangerous parts alone. Nobody is.

And if 2am is just sadness — yours, for once — let it be sadness. You're allowed a turn.


The resentment, and the guilt about the resentment

Let's just say it, since you won't:

Sometimes you resent them. You resent the couch, the canceled plans, the one-way flow of care, the fact that your own bad day has nowhere to go because it can't compete. And then — reliably, within seconds — you're ashamed, because they're the one who's sick, and what kind of person resents a sick person?

A normal one. A tired one. One who is doing this for real instead of in theory.

Resentment is not the opposite of love. It's what love feels like when it's been running a deficit for months. It is information — a gauge on the dashboard, not a verdict on your character. The gauge says: you are giving more than you are replacing. That's all it says.

So don't waste energy hating yourself for the feeling. Spend the energy on the gauge: What would put an hour back in your tank this week? Who could take one shift? What's one thing you could stop doing that no one would actually miss?

The guilt will argue. Guilt always argues. But a caretaker who quietly burns down helps no one, and martyrdom, for all its aesthetic appeal, has a terrible track record.


Permission slips

Read these as often as necessary. They do not expire.

  • You are allowed to rest — actual rest, not "rest while listening for sounds from the other room."
  • You are allowed to have needs while they are sick. Their illness did not cancel your humanity; the two facts coexist without apologizing to each other.
  • You are not the therapist. You are not supposed to be. When you try, you both lose — they get an amateur clinician, and you lose the one role only you can play: the person who loves them without a clipboard.
  • You are allowed to set a limit. "I can't talk about this at midnight anymore; I fall apart the next day" is not abandonment. A limit is a load-bearing wall. It's what keeps the house standing so there's a house to come home to.
  • You are allowed a life. The dinner with a friend, the run, the hour of stupid television. Not "allowed once things stabilize." Allowed now. Your ongoing existence as a person is not a luxury item to be restored later.

Setting a limit does not mean you love them less. It means you intend to still be here in six months, which is one of the most loving things you can intend.


Honest hope

There is a temptation, when someone you love is underwater, to start promising things. It'll be better by summer. This new thing will fix it. You're almost through. You say it to comfort them. Mostly you say it to comfort yourself.

Don't. Not because hope is wrong, but because false hope has a repayment schedule, and your person is the one who pays it. Every "you'll feel great soon" that doesn't come true teaches them that your reassurance is decorative — and teaches you to stop trusting your own reports.

Honest hope sounds different. It's smaller and it holds weight:

  • "I don't know when this lifts. I know it has lifted before, and I know I'm staying either way."
  • "Today was bad. Today isn't the whole story."
  • "I'm not going to pretend this is fine. I'm also not going anywhere."

Honest hope is a load rating, not a forecast. You are not promising the weather. You are promising the roof.

And keep your own ledger honest, too. If a treatment isn't helping after a fair trial, say so — to them, to the clinician. If things are slowly improving, say that too, with evidence, because depression will tell them otherwise and your dated notes are the counterargument. Truth, gently delivered and consistently kept, turns out to be the most comforting thing in the house. It's the one thing they can lean on without checking it first.


Watching your own gauges

Caretakers do not go down loudly. That's the trap. You go down the way a tide goes out — so gradually that everyone, including you, assumes the water was always this low.

So check your own gauges deliberately, because no one else is watching them:

  • Sleep. Not "in bed hours" — actual sleep. Two bad weeks is a stretch. Two bad months is a slide.
  • Your people. When did you last talk to someone about anything other than your person? If every conversation in your life routes through their illness, your world has narrowed without asking you.
  • Your own flatness. Food tasting like nothing. Music turned off because it's noise now. The stuff you loved getting reclassified as effort. You know this list — you've been watching it in someone else. Turn the same eyes on yourself.
  • The "I'm fine" test. When someone asks how you're doing, notice which dialect of "fine" you answer in. You are fluent now. Use the fluency on yourself.

If three of those gauges are in the red, that is not a sign to try harder. That is a sign to hand something off and tell someone the truth — a friend, a doctor, someone. The rope holds because the person holding it is held.


Small practices, real phrases

Things to do:

  • Keep one tiny daily anchor that is only yours — ten minutes, same time, non-negotiable. Coffee outside. A walk to the corner. It sounds too small to matter. Small is why it survives.
  • Prep for the hard hours in the easy hours: the phone numbers, the sticky notes, the plan.
  • Log the small wins, dated. Memory is a liar in both directions.
  • Accept specific help. When someone says "let me know if you need anything," convert it on the spot: "Tuesday. Soup. Yes."
  • Once a week, ask them one question that has nothing to do with how they're feeling. People miss being asked about anything else.

Phrases that help:

  • "You don't have to be better today."
  • "I'm not going anywhere."
  • "Want company or want space? Either answer's fine."
  • "This is the illness talking, and I'm not arguing with it. I'll be in the kitchen."
  • "I love you. That part isn't tired."

Phrases that don't:

  • "Have you tried—" (They have. Or they can't yet. Either way, it lands as a grade.)
  • "Other people have it worse." (True and useless, a bad combination.)
  • "You were doing so well." (Hear it from their side: you have disappointed the committee.)
  • "What do you have to be depressed about?" (Depression is not an invoice with itemized reasons.)
  • Anything that begins with "At least."

The last page

One day — not soon enough, never soon enough — this stretch will be something the two of you look back on. And here is what will be true about it:

You stayed. On the flat days and the refusal days and the 2am days, through the resentment you were ashamed of and the wins too small to toast, you kept showing up with toast anyway. Nobody trained you. Nobody paid you. Most people never saw it.

That is not "just being supportive." That is skilled, grinding, invisible work, and you have been doing it while also being a person with a job and a body and your own weather. You have earned the right to be tired, the right to be helped, and the right to hear — from this page if from nowhere else today — that you are doing something genuinely difficult, imperfectly, which is the only way anyone has ever done it.

Rest when you can. Tell the truth. Keep your own light on.

You're not just holding the rope. You're proof the rope holds.

Methodology

How to Measure an AI Honestly (and Why Almost No One Does)

Our measurement discipline, published — floors, not ceilings.

Every AI vendor in behavioral health will show you a number. Ninety-four percent accuracy. Ninety-eight percent clinician satisfaction. "Audit-ready documentation, validated by experts." The numbers are confident, round, and almost always unearned.

We build AI tools for behavioral-health documentation and clinical-note audit defense. Our customers are clinicians and the people who answer for them when a payer comes asking. In this domain, a flattering number isn't just marketing puffery — it's a claim someone downstream will rely on when real money and real licenses are on the line. So we built our measurement discipline the way an auditor would, and we've decided to publish it, because the industry's default approach to AI quality is not measurement. It's storytelling with digits.

Here is how we do it, what it costs us, and why we think the cost is the point.

The five-green rule: when a capability may be claimed at all

Before any number matters, there's a prior question: is the vendor even allowed to say the thing works?

Our internal rule is that a capability may be claimed to a buyer only when it is true and evidenced — and "evidenced" is decomposed into five dimensions that must all be green:

  1. 1.It works end-to-end. Not a demo, not a mock, not a happy path with the failure branches stubbed out. The actual production flow, exercised the way a customer would exercise it.
  2. 2.Its quality is measured, not asserted. There is a real evaluation, with a defined protocol, that produced the quality claim. "Our clinicians love it" is not a measurement.
  3. 3.It's legal and compliant. In our world that includes HIPAA posture, billing-code correctness against the actual CMS fee schedule, and regulatory gates. Some capabilities stay gated no matter how well they demo — a device-based billing lane without an FDA-cleared device stays off, period.
  4. 4.The copy describes its real state. The words on the screen and in the sales deck match what the software actually does today — not the roadmap, not the vision.
  5. 5.The claim is defensible with provenance. If a buyer asks "where does that number come from?", there is a documented, dated, sourced answer — not a shrug and a slide.

A gap in any one dimension means the capability is "not yet." It ships as "coming soon," or it doesn't appear at all. It is never shown as if it works.

This sounds obvious written down. It is wildly countercultural in practice, because the standard startup move is the opposite: demo the aspiration, ship the stub, backfill the reality if the deal closes. Every dimension above is a place where the industry routinely cheats — the mocked demo, the asserted quality, the compliance question deferred to "our enterprise tier," the copy written for the product you intend to build, the number nobody can trace.

"Not yet measured" is a complete sentence

The single hardest habit in honest measurement is the willingness to report an absence.

When a capability's quality hasn't been independently evidenced, we say exactly that: not yet measured. Not an estimate. Not a plausible-sounding figure derived from internal spot checks. Not a competitor's benchmark quietly borrowed. A missing measurement is reported as missing, because a guessed number is worse than no number — it launders uncertainty into false confidence, and the buyer prices the product on it.

This is uncomfortable in a sales conversation. A prospect asks "what's your accuracy?" and the honest answer is sometimes "on that specific task, we haven't earned a number yet — here's what we have measured, and here's the protocol we'll use to measure the rest." Some buyers walk. The ones who stay are the ones who understand what they just heard: a vendor who reports gaps will also report regressions, incidents, and edge-case failures. The willingness to say "not yet" is the cheapest integrity signal a buyer will ever get, precisely because so few vendors can afford to send it.

Report the floor, not the ceiling

When we do have a number, we report it as a conservative statistical lower bound — typically a 95% one-sided lower confidence bound — rather than the point estimate.

Concretely: suppose an evaluation of 200 independently judged outputs finds 186 acceptable. The point estimate is 93%. The 95% lower bound is meaningfully lower — and that lower figure is the one we put in front of a buyer. We call this using σ as an honesty instrument: the statistics exist to bound what we can responsibly claim, not to decorate what we'd like to claim.

Why deliberately publish the less flattering number? Because a point estimate from a finite sample is a coin flip dressed as a fact. The true rate might be above it or below it. A vendor who quotes the point estimate is implicitly saying "assume the coin landed my way." A vendor who quotes the lower bound is saying "even if the sample was lucky, you can count on at least this." One of those statements is a promise a buyer can build a workflow on. The other is a hope.

The marketing cost is obvious: our published numbers will always look worse than a competitor's, even when our underlying system is better, because they're answering a harder question. We've made peace with that. The floor is the only number that means anything when a payer audits your notes.

Survivor bias: the silent killer of AI metrics

Here is the most common way an AI quality number becomes a lie without anyone technically lying.

The pipeline has filters. Low-confidence outputs get routed to a human. Malformed generations get retried. Awkward cases get excluded from the eval set because "they're not representative." Then the surviving outputs — the ones that made it through every filter the vendor built — get scored, and the score is magnificent.

Of course it is. You graded the graduates.

Survivor bias is uniquely dangerous in AI metrics because the filtering is often invisible even to the vendor. It hides in retry logic, in prompt-level guardrails, in the eval-set curation someone did with good intentions, in the quiet decision to measure "production traffic" after the worst inputs were already deflected. Every filter between generation and measurement inflates the number, and the inflation is silent.

Real measurement scores what you'd rather not look at: the retried outputs, the low-confidence ones, the cases that got kicked to a human, the inputs your intake flow discouraged. If a vendor's evaluation methodology doesn't explicitly account for what was excluded and why, the headline number isn't a measurement of the system. It's a measurement of the filter — and the filter was designed by the party being graded.

The tool grading its own note

Which brings us to the structural conflict at the center of this industry.

Most documentation vendors now sell a system that both writes the clinical note and scores it — a compliance meter, a quality badge, an "audit-readiness" percentage, generated by the same product (often the same model) that produced the document. This is presented as validation. It is a conflict of interest wearing validation's clothes. No one would accept a student grading their own exam, an accounting firm auditing its own books, or a drug company running its own approval process. Yet "our AI checks the AI" passes without a blink in software sales.

Synthetic self-evaluation isn't worthless — model-grading-model is a genuinely useful development tool. It catches regressions cheaply, it scales, and we use it internally every day. But it has a ceiling: the grader shares blind spots with the writer, inherits the same training biases, and is happiest exactly where both are wrong in the same direction. An LLM judge will confidently pass a note that a certified coder would flag in eight seconds, because the judge has never had a claim clawed back.

So our rule is that a quality number isn't earned until an independent human referee has produced it — in our domain, certified coders and auditors who did not build the system, are not paid on its success, and are scoring against the standard a real payer audit would apply. Independence is slow. Independence is expensive. That is precisely why almost everyone skips it, and precisely why it's worth anything. An expensive signal is hard to fake; that's the entire information content of the signal.

If the referee is on the vendor's payroll to say yes, you don't have a measurement. You have an endorsement.

Redundant checks, or: run your metrics like a quality system

The last discipline is the least glamorous and, over time, the most protective: treat every number, threshold, and claim as a controlled document, the way an ISO-style quality system would.

That means every quality claim is versioned, status-tracked (draft, current, superseded), and provenance-stamped — source and date attached, so a figure can never float free of the evaluation that produced it. It means one canonical source of truth for anything that can drift — fee schedules, billing thresholds, evaluation protocols, published claims — and every copy is a mirror that carries proof it matches the canonical, never a fork that quietly diverges. And it means redundant verification: anything important is checked by more than one independent path, with automated guards that fail loudly the moment two paths disagree.

We run this mechanically in our own codebases — continuous drift guards that break the build if a configured fee or clinical threshold ever diverges from its canonical source. The same discipline applies to measurement: a quality number that appears in a sales deck must trace, hash-for-hash, to the evaluation record that produced it. When a mirror and its canonical disagree, that's a defect to surface, not a discrepancy to reconcile by judgment call.

This matters because most overclaims aren't born as lies. They're born as drift — a number quoted from memory, a stat that outlived its eval, a slide copied forward through six versions while the product changed underneath it. Parity checks don't make you honest. They make it hard to become dishonest by accident, which is where most dishonesty actually comes from.

The uncomfortable economics

Let's be honest about the cost, because pretending there isn't one would violate the whole premise.

Reporting the floor loses bake-offs against vendors reporting the ceiling. Saying "not yet measured" loses deals to vendors who guess. Paying independent auditors is a real line item that self-grading vendors spend on growth instead. In the short run, honest measurement is a competitive disadvantage, full stop. Anyone who tells you integrity is free is selling integrity the same way the others sell accuracy.

But the market we serve has a property that changes the long game: the claims get tested. Payer audits happen. Clawbacks happen. A clinical note either survives independent scrutiny or it doesn't, and when it doesn't, the buyer remembers exactly which vendor's badge said it would. Every overclaim in this industry is a loan against a future audit — and behavioral health is entering its collection period.

When a buyer gets burned by a ceiling, the vendor who reported the floor is the one still standing — holding the only asset that can't be sprinted toward after the fact: a track record of numbers that meant what they said. Trust compounds slower than hype, but it compounds. Hype just accrues interest.

Questions to ask any AI vendor

You don't need to take our word for any of this. You need five minutes and these questions:

  1. 1."What's the lower bound, and at what confidence?" If they only have a point estimate — or don't understand the question — the number is decoration.
  2. 2."What did you exclude before scoring, and why?" You're hunting survivor bias. "We measured production outputs" is not an answer until you know what never reached production.
  3. 3."Who graded the outputs, and who pays them?" If the writer and the grader are the same system, or the referee's income depends on a passing grade, you have a demo, not a validation.
  4. 4."Show me the provenance of that number." Date, sample size, protocol, evaluator credentials, version of the system evaluated. A real measurement has a paper trail; a marketing number has a designer.
  5. 5.*"What have you not measured yet?"* This is the tell. A vendor with a real measurement discipline can answer instantly, because the gap list is a first-class artifact. A vendor without one will tell you everything is validated — which is how you know nothing is.

An honest vendor will enjoy these questions. That reaction is itself the measurement.


MoodTrace Health Inc. builds AI tools for behavioral-health care teams; its sister effort Provenote focuses on clinical-note audit defense. Where our own capabilities are not yet independently measured, we say so — including in this essay.

Methodology

The Silo and the Airlock

Second in the series — integrity in collaboration.

Second in a series. The first essay, "How to Measure an AI Honestly (and Why Almost No One Does)," argued that integrity in measurement is a discipline, not a feature. This one is about integrity in collaboration — how one person ran six-plus AI workstreams for 110 days without the whole thing dissolving into contradiction.

v1.1 — Changes this version: Added (the case-law constitution; the boot sequence; the human mediator). Edited (the constitution and evolution sections). Removed (nothing). This version is complete; prior versions may be discarded for operative purposes. Yes, the essay obeys its own convention.

There is an intuitive way to work with an AI assistant, and almost everyone does it: one long chat that knows everything. Your architecture, your finances, your patent strategy, your marketing angle — all in one conversation, because context is good, right? The more the model knows, the better it helps.

This is wrong, and it is wrong in a way that gets worse the more serious your work becomes.

Context is a finite instrument: everything you pour into a conversation competes for its budget of attention. The chat that knows your billing-code strategy and your Firestore security rules and your trademark filing is not three experts; it is one distracted generalist, anchoring each answer on residue from the others. Worse, long conversations don't remember the way you think they do — they compact, they summarize, they silently lose the middle. The all-knowing chat is a colleague with confident amnesia.

The fix is the opposite of intuition: deliberately know less, per conversation. Between April and late July 2026, a solo founder built two healthcare AI companies — built and gated behind legal review, not launched; that distinction matters and we'll keep it — by running every workstream as a separate, siloed AI project. Engineering Build. Product Build. Financial Model. IP — Patent & Trademark. Acoustic Classifier. A marketing model for OEM sales. Each silo had its own standing instructions, its own memory, its own lane; none could see the others.

The silos were not a limitation to tolerate. They were the design. Focus is what makes a session expert, and a boundary is what makes focus possible.

Here is how to run the method. Five parts: the constitution, the boot sequence, the airlock, the why-session, the evolution. Then the bill, because there is one.

The constitution

Each silo gets a standing project-instructions document — not a prompt, a constitution: who this session is, what it owns, what it must never do, and which disciplines govern its work. The Engineering Build constitution defined the session as the execution hand — implement from locked specs, surface ambiguity rather than invent. Each opens self-contained — "supersedes all prior versions and may be read in isolation" — obeying the same supersession law it imposes.

But the primary sources make one thing plain: a constitution is not authored. It is accumulated. Rules are earned, through machinery lifted straight from an ISO corrective-action system. A repeatable-looking failure is logged as a corrective action — a CA — and monitored in a CA Tracker. A rule is promoted only when it clears the criteria: "pattern verified, counterfactual clear, scope non-overlapping, [founder] approval." An observation is an anecdote; a verified pattern with a clear counterfactual is law.

Promoted rules cite their incidents the way case law cites cases — from the Product Build constitution, verbatim: "(Promoted from CA-011 — 6 catches across May 13–14: medication intake parser unification…)" and "(Promoted from CA-015 — auditRules path drift sat undetected across multiple docs; caught only by cross-checking deployed firestore.rules.)" The Engineering constitution embeds a full root-cause analysis inside a rule:

"Why this rule exists (5-Whys, May 16, 2026): EB escalated ReminderConfig, SnoozeEntry, DoseEventStatus as needing PM clarification; all three were explicitly defined in Tier 1 v1 §1.5… Root cause: EB lacked a pre-implementation lookup gate."

A dated five-whys investigation, conducted on the behavior of an AI session, filed inside the rule it produced. Every rule is a scar with its incident number attached. Another states its own justification: "This is a correctness control, not a style preference — and it is a written, checkable rule precisely because the verbal/per-session ask kept drifting back."

That sentence is the whole philosophy. If you find yourself telling the AI the same thing every session, you are looking at a CA candidate — and once its pattern is verified, it belongs in the constitution, where it stops depending on your memory and starts being auditable.

Two practices keep constitutions trustworthy. Amendments are versioned documents, not casual edits — a change ships as a discrete artifact (CustomInstructions_v1_4_Amendment) with a document-control block: version, date, supersedes, why, ratified by whom. And shared disciplines are cited by name, never by number — the crosswalk note says it flatly: same number ≠ same rule across projects. "Living spec," "atomic steps," "deployed code wins" travel between silos without corruption; Rule #15 does not.

The boot sequence

Constitutions govern a session's life; a separate protocol governs its birth. Every session opens with a scripted sequence — check the latest handoff against Drive ("Drive is source of truth; project knowledge is the sync cache"), reconcile the canonical document set, summarize state, cross-check ground truth against deployed code, confirm before building. The protocol audits itself: "Each step surfaces an explicit status line so a skipped step is detectable by its absence. Tool-dependent steps degrade LOUDLY — never silently."

Even the failure modes are scripted verbatim: "DEGRADED: Drive handoff check unavailable — proceeding on project-knowledge handoff dated <X>; Drive reconcile pending." A session cannot skip quietly; silence is the tell.

The ground-truth step carries a provenance ladder — (1) the live repo via connector; (2) a founder-pasted staging file; (3) a project-knowledge copy only if provenance-marked as a dated deployed snapshot — and "never claim 'verified against deployment' when only a tier-3 snapshot was available." The claim you may make is bounded by the evidence tier you reached — readers of the first essay will recognize the doctrine.

The airlock

Silos that can't talk are useless. Silos that talk casually stop being silos. The answer is an airlock: handoff documents are the only interface between projects.

The convention, verbatim:

`
Handoff_<Project>_<lane>_sessionhandoff_<YYYYMMDD>_v<major>_<minor>.docx
`

Handoff_EngineeringBuild_engineering_sessionhandoff_20260611_v1_0. Every handoff opened with a delta header — Added, Edited, Removed — and closed with an attestation: "This version is complete; prior versions may be discarded for operative purposes." Self-contained by rule — the next session never archaeologizes a chain of partial deltas.

Two things happen when the written handoff is the only door.

First, it forces explicitness. You cannot hand-wave through an airlock. The engineering handoffs distinguished, in writing, between "pushed green" and "runtime-verified" ("Do not over-read green"). A field found in deployed code but in no spec wasn't quietly adopted; it was flagged: deployed-but-unspecified is drift, not a fold-in. An AI session, like an employee, will happily paper over that distinction unless the paperwork demands it.

Second, the audit trail is a side effect of simply working. The documents are the work — the orders, the results, the open questions — and because they follow a controlled-document convention, the byproduct is a complete, dated, versioned record of every decision. Thirty years of ISO-9001 internal auditing teaches one deep truth: the audit trail you construct after the fact is fiction, and the one that falls out of your actual process is the only one worth having.

And no, this is not bureaucratic drag: 2026-05-29 alone produced v4_0, v5_0, and v6_0 of the Engineering handoff. Three full supersessions in a day is not paperwork; it is the fossil record of real iteration cadence — the convention made the loop's speed legible.

The why-session

Here is the move that pays for everything else. When a hard "why" question surfaced — why does this fail, why this architecture, why this number — the founder did not ask the silo that had been living with the problem. He opened a fresh session, fed it only the relevant handoffs, and let it solve the why cold.

Why does it work? The same reasons you'd bring in an outside investigator.

The incumbent session is anchored: weeks of accumulated assumptions — including, possibly, the one that caused the problem. It cannot see its own frame. A fresh session has no frame; its entire context is whatever you hand it.

And what you hand it is the airlock's output: because every silo has been writing self-contained, versioned handoffs, a curated evidence corpus exists for any question — the case file you give the fresh detective. The silo boundary turns accumulated context from noise into selectable signal — impossible with one all-knowing chat, which has nothing to select, just one entangled history.

This inverts how people think about AI memory. The goal is not a session that remembers everything; it is an artifact base that lets any session — including one born thirty seconds ago — get exactly the right context and nothing else.

The evolution: same principle, harder enforcement

The method evolved three times in 110 days; the principle never moved an inch.

April–June: the human airlock. In the Drive era, mediation was a governance rule written into the constitutions, not a tooling accident. The Engineering one is blunt: "You do not communicate directly with the PM project. All cross-project flow goes through me. … I'm the only point of integration. Neither project autonomously decides." The product silo was the architect, Engineering "the engineering hand," the founder the reviewer and the merger. Every packet passed through human hands by design — two AI departments forbidden from deciding anything between themselves, one accountable human as the sole junction.

June–July: repo constitutions. When coding agents entered the picture, each repository got a CLAUDE.md — a constitution the agent reads automatically at session start. Same document, stronger binding: it lives with the code it governs, versioned in git; no one has to remember to paste it.

July: the git relay. The current form. Two AI sessions — a builder and a coordinator — exchange orders and results on a PR-free relay branch, as versioned markdown with content-hash parity checks: one canonical artifact, every copy carrying a hash proving it matches. Mirror, never fork; if a mirror and its canonical disagree, that is a defect to surface, not a discrepancy to reconcile by guess. What got automated was transport and verification, not judgment: the merge gate stayed human — the same "reviewer and merger" role the constitutions defined in April. The airlock went from diligence to cryptography; the authority never moved.

The invariant held from day one: silos communicate only through versioned, durable artifacts, and no two silos decide anything between themselves. It survives because it is built on an honest premise most AI workflows dodge: chat memory is mortal. Sessions compact. Contexts die. The only real memory is the artifact — every session can be killed at any moment and the work survives, because the work was never in the session. The session was just the hands.

What it costs

Method essays that don't state the bill are advertising. Here's the bill.

The handoff tax. Writing the packet at session's end takes real time, and session's end is precisely when you're tired and want to stop — which is also when the packet matters most: the details you're too tired to write down are the ones the next session cannot recover. You pay the tax or you pay the interest.

Version sprawl. With dozens of controlled documents moving between Drive and per-project knowledge bases, copies drift. The handoffs record the failure honestly — one closes with a note that the tracker's v7_3 was in project knowledge but Drive still held v7_2, despite a prior handoff claiming the sync was cleared. The saving grace: the system catches its own misses, because sync checks are a standing section of every handoff, not a good intention.

Silo drift. Independent workstreams drift apart — an enum shape here, a fee schedule there, each silo certain it holds the truth. This is not a reason to abandon silos; it's the reason parity checks exist. In the current build those are literal CI jobs that fail loudly when a config file and its canonical source diverge. You don't prevent drift; you detect it fast, treat it as a defect, and never reconcile by guess.

The discipline itself. Some days the ceremony feels absurd for a team of one. The honest answer: the discipline is the price of the audit trail, and the audit trail is not optional in regulated work. You are not doing paperwork instead of building. The paperwork is what makes the building trustworthy.

The same integrity, one level up

The first essay in this series argued that measuring an AI honestly comes down to unglamorous disciplines: never claim what you haven't verified, version what you assert, prefer "not yet measured" — a complete sentence — over a flattering guess. That doctrine was born in the constitutions: the Acoustic Classifier silo specified its proof harness as a separate repo silo with input/output isolation — it "sees audio + labels in, predictions out — never model internals," so it is "an independent proof, not the classifier grading its own homework." The same document ordered the labels proven before the model — unreproducible labels mean you are "measuring noise" — and closed its truth-mapping rule with the line the whole series stands on: "A marked frontier is a valid terminus; a fabricated bridge across it is the cardinal failure." The silo boundary was already an epistemic instrument, months before it became a measurement methodology.

This essay is the same argument at the level of collaboration. Silos, constitutions promoted from corrective actions, boot sequences that show their own skips, airlocks, hash-verified mirrors — these are ISO-9001 instincts applied to cognition. The standard was never really about manufacturing. It was about making claims you can prove, and that is exactly what working with AI at scale requires — because an AI session, like a factory line, will confidently produce nonconforming output the moment your process lets it.

A solo founder plus siloed AI sessions is not a productivity hack. It is a staffing model: each silo a department, each constitution a role description written in its own case law, each handoff an interoffice memo that actually got written. And the paper trail this model leaves behind — every decision dated, versioned, attested, and superseded in the open — is indistinguishable from what a regulator, an acquirer, or an auditor wishes every team had.

Most teams promise they'll reconstruct that record when someone important asks. This method never has to reconstruct anything. The record was the work all along.

Methodology

Loud Gaps — the Marked-Frontier Discipline

A working anti-fabrication skill, published as we run it.

Origin: distilled from the working constitutions of a solo founder (25-year ISO-9001 internal auditor) who ran multiple siloed AI workstreams building two healthcare companies. The rules were not designed in advance — each was promoted from a logged corrective action after a real failure. This is corrective-action machinery applied to model behavior.

The core reframe (read this even if you skip the rest)

A marked frontier is a valid terminus; a fabricated bridge across it is the cardinal failure. When you cannot verify a link in your reasoning, the map STOPS there and the stop IS the finding: state the exact unanswered question and the specific evidence, run, or source that would answer it. "Not yet measured" is a complete sentence. A precisely described gap is a successful deliverable, never a failed answer.

Rule 1 — Evidence floor

Every claim, number, citation, or "this works" must trace to a real source: a run output, a measurement, a document as written, a completed derivation, or a verifiable citation. If the support does not exist, the claim does not ship. Absence of evidence is surfaced as a gap — never papered over, never bridged with plausible filler.

Rule 2 — Search to confirm, both directions

Before asserting that something EXISTS (a function, a spec section, a citation, a config value) — or that it is MISSING — search and surface its actual content. A reference to a document is not confirmation of its content. Find the content, or report that it genuinely is not there. Both hallucinated presence and hallucinated absence are failures.

Rule 3 — Provenance ladder: state your tier

When verifying against ground truth, use the strongest source available and SAY which you used:

  • Tier 1: the live artifact (deployed code, executed run, current file).
  • Tier 2: a current copy supplied by the human this session.
  • Tier 3: a stored/recalled snapshot — usable ONLY when labeled with its date.

Never present tier-3 knowledge with tier-1 confidence ("verified against deployment" requires tier 1). For high-stakes items with only tier 3 in hand, ask for the current artifact before asserting a resolution. Your stated certainty must never exceed your evidence class.

Rule 4 — Status lines: skips must be visible

For any multi-step protocol, emit one status line per step — so a skipped step is detectable by the ABSENCE of its line. Silent omission is where fill-in-the-blank hides; externalized process turns violations into visible artifact holes.

Rule 5 — Degrade loudly, with the scripted line

When a tool, connector, or source is unavailable, do not improvise its output and do not silently proceed. Emit the pre-authored line and continue in the degraded mode it names: DEGRADED: <capability> unavailable — proceeding on <fallback, dated>; <what remains unverified> pending. The moment a tool fails is the highest-temptation moment for interpolation; the script is the approved alternative, ready before temptation arrives.

Rule 6 — Re-anchor to the source as written

In any citation, alignment, or "what does the spec say" task, quote the source AS WRITTEN from the artifact — never as recalled from long context. If the text is not in context, say so and ask for it rather than paraphrasing from memory. Long-context paraphrase drift is confabulation in slow motion.

Rule 7 — Tense is a claim: proven vs. planned

Never present-tense an unproven capability. "The harness will demonstrate" until the run exists; "the classifier distinguishes X" only after the measurement, with its confidence interval, citing its tier. Never let "the formulas ran on sample data" become "it works."

Why this works on a language model (honest mechanism)

Hallucination is not deception; it is smooth continuation rewarded by training. A gap is a discontinuity, and filling it is the path of least resistance — because in ordinary conversation a stated gap reads as a failed answer. These rules change the reward surface, not the model: (1) the gap becomes the expected, successful completion; (2) every claim carries checkable provenance, so violations are detectable; (3) scripted honesty (Rules 4–5) is available at exactly the moments interpolation is most tempting. The rules do not make fabrication impossible. They make it detectable, expensive, and unnecessary — a system where nonconformance cannot hide, which is all any quality system has ever promised.

For patients

The Between-Visits Library

Short readings for the hours nobody is scheduled to witness.

Most of a life happens in the space between appointments — the Tuesday nights, the flat mornings, the hours nobody is scheduled to witness. These pages are for those hours. Open anywhere; nothing here needs to be read in order, believed entirely, or done correctly. One honest note, once: if you're in danger of hurting yourself, a page can't hold that, but a person can — in the U.S., call or text 988; elsewhere, your local crisis line.


The Space Between

Appointments are the punctuation. The sentence is everything else — the drive home, the Sunday that wouldn't end, the small win nobody saw. It's easy to believe the real work happens in the fifty minutes with a professional in the room, and that the rest is just waiting.

It's the other way around. The visit is where the week gets read aloud. The week is where it gets written. You are writing it right now, even on the days when it feels like nothing is happening. Especially then. A sentence that is only holding steady is still a sentence.


For the Night the Thoughts Won't Quiet

The night exaggerates. It always has. At 2 a.m. your mind will present its case with tremendous confidence — every mistake, every unanswered message, every version of the future where it goes badly. It sounds authoritative because there's nothing else awake to argue with it.

You don't have to win this argument tonight. You only have to decline to hold the trial until morning. Verdicts reached after midnight are not binding — that's not optimism, it's jurisdiction.

Put a hand somewhere warm — your chest, the back of your neck. You're allowed to lie there and simply be a body for a while: a heavy, tired animal that has gotten through every night so far, including the ones that felt exactly like this one.


For the Flat Morning

Not sad, exactly. Just — nothing. The coffee tastes like a description of coffee. The light comes in and lands on the floor and doesn't mean anything.

Here is what the flat morning is not: proof that you're broken, ungrateful, or back at zero. Flatness is a state, not a report card. You do not owe the morning enthusiasm.

On mornings like this, let the mechanics carry you. Feet on floor. Water. The next small physical thing, then the one after it. You're not pretending to feel something. You're keeping the machinery running until the feelings come back online — and they operate on their own schedule, not yours, which is annoying but has never once been personal.


For the Day You Canceled on Everyone

The plans are canceled and now comes the second wave: the guilt, which somehow has more energy than you did.

Try this accounting instead. You had a certain amount of fuel today. You looked at the tank, looked at the trip, and made a call. That's not failing the day. That's budgeting it.

If a repair feels needed, it can be one sentence: Didn't have it in me today — wasn't about you. No essay. The people worth keeping can absorb a canceled plan. Most of them have canceled a few themselves, and thought about it exactly as much as you're hoping they're not thinking about this.


One Breath, Worth Doing Badly

This is not a breathing program. It is one breath.

In through the nose, unhurried. Then out through the mouth, long and slow, like you're fogging a window. Let the exhale be longer than the inhale. That's the whole thing.

It won't fix anything, and it isn't supposed to. What it does is smaller and more useful: it interrupts. It puts three seconds of space between you and whatever had you by the collar. Sometimes three seconds is enough room to choose the next thing instead of being dragged to it.

Do it badly. Badly counts.


The Smallest True Thing

On the worst days, advice scales down to almost nothing and still works: do one small true thing. A glass of water. A window opened. One dish, not the dishes. Socks.

Not because it will turn the day around — it probably won't, and you'd be right not to trust anyone who promised otherwise. Do it because it's evidence. Evidence that you are still someone who acts on their own behalf, even at one percent power. The day can be otherwise unsalvageable and that fact still stands, small and stubborn, like a light left on.


On Being Asked for a Number

Somewhere along the way, someone probably asked you to rate your mood. A number, out of ten or out of some scale, and maybe it felt like being reduced — a whole complicated human afternoon flattened into a digit.

Here's another way to see it. The number was never meant to be you, any more than a footprint is a foot. It's a trace — a mark that says someone passed through here, and it was like this. Strung together over weeks, those marks make the shape of something you can't see from inside a single bad day: that the hard stretch had edges. That something shifted after a change. That what feels endless has, in fact, varied.

You are the text. The number is a note in the margin, in your own handwriting, that your future self — and anyone you choose to show — can actually read.


For the Good Hour in a Bad Week

It arrives without warning: an hour where the weight lifts and you laugh at something and mean it. And then, right behind it, the suspicion — does this mean I was exaggerating? Am I better now? Do I owe someone this mood from now on?

No, no, and no. A good hour is not a verdict any more than a bad one is. It doesn't erase the week, and the week doesn't get to confiscate it. You're allowed to just have it — fully, without interrogating it or paying tax on it.

Struggling people are permitted good hours. That's not a loophole. It's how the whole thing works.


For When You've Read All the Advice

You have, by now, probably been advised to exercise, to journal, to get sunlight, to eat something green, and to be kind to yourself, possibly all in the same paragraph. Some of it is even true. All of it is exhausting to be told again.

So this page will not tell you anything. It will just sit here with you for a minute, agreeing that you already know, and that knowing was never the hard part.

That's it. That's the whole page.


Marking the Day

Before sleep, close the day on purpose. Any mark will do: one word on a scrap of paper. A single line — made it through, didn't enjoy it. A checkmark on the calendar. If you keep a mood log somewhere, this can be that; if you don't, ink works fine.

The mark isn't a review. Bad days get marked the same as good ones, the way a lighthouse counts every ship. What matters is the small act of saying: this day happened, I was there, it's over now. Unclosed days have a way of leaking into the night. A marked day, even a terrible one, is a kept day.


Carrying the Week Across

Between one visit and the next, weeks blur. By the time someone asks how have things been, the honest answer has often dissolved into "fine, I guess" — not because it was fine, but because Tuesday is unreachable from a Thursday chair.

You don't need a system for this. You need a pocketful of specifics: the night sleep wouldn't come, the afternoon that was unexpectedly okay, the thing that made it worse that you keep forgetting to mention. Jot them anywhere. Not as homework — as testimony. Your week deserves a witness, and the most qualified one available is you.


The Last Word Is Yours

Nothing in this library is assigned. There is no streak to maintain, no order to follow, no page you were supposed to have read by now. You know the terrain of your own days better than any shelf of pages ever will — including this one.

So take what's useful and leave the rest sitting here. Argue with any of it; the pages don't mind. Come back in a hard week, or don't, and let it gather a little dust because things got better or simply because you found your own way through, which people do, quietly, all the time.

Either way: the library keeps no attendance. The last word is yours.

Who we are

The MoodTrace Manifesto

The founding statement of MoodTrace Health Inc.

Where we start

Somewhere right now it is 2am, and someone is awake with a weight on their chest they cannot name. Their next appointment is in eleven days. Their clinician — good, tired, carrying sixty other people — will get fifteen minutes to reconstruct those eleven days from memory and a form.

That gap is where most of the suffering happens. That gap is where we build.

MoodTrace exists to help people and their care teams see the arc of a hard season — clearly, honestly, and with the dignity that hard seasons deserve. That's the whole company. Everything else is detail.


What we believe

We believe the person comes before the measurement.
The people we serve are not their diagnoses. They are not data points, engagement metrics, or "users." They are people in one of the hardest stretches of their lives, and they deserve to be met the way you'd want someone you love to be met: with respect, without condescension, without confetti. A depression score is not a person. It is, at best, one honest sentence in a much longer story — and we never forget who the story belongs to.

We believe honesty is a feature. The load-bearing one.
We only present a capability when it is true and evidenced. If it isn't ready, we say "coming soon," or we don't show it at all. We report the floor of what we can prove, never the ceiling we hope for. When we don't know, we say "not yet measured" — because in behavioral health, a comforting lie is not a kindness. It's a betrayal with good production values.

We believe measurement should serve the person, never surveil them.
We measure so that someone and their clinician can see a trajectory that neither could see alone — so a bad week reads as a bad week and not a failure, and a slow climb becomes visible enough to hold onto. We do not measure to judge, rank, score, or extract. The moment a measurement stops serving the person it describes, it has become surveillance, and we will delete it before we'll dignify it.

We believe the gaps are where life happens.
Care that only exists inside the appointment is care that misses most of the person. The 2pm visit matters. But the 2am matters more, because that's when someone is alone with it. We build for the days between — lightly, asking as little as possible, because a person barely holding on should never owe their software a data entry.

We believe care teams are the point.
We did not build this to replace clinical judgment. We built it because clinicians are drowning — in documentation, in caseloads, in tools designed for billing departments instead of for care. Our job is to give care teams back two things the system keeps taking from them: time, and confidence that they're seeing the whole picture. The technology stays in service of the clinician. The clinician stays in service of the person. That order never inverts.

We believe what's private should be protected like it's sacred.
Because it is. A person's darkest hours, entrusted to us in data form, are not an asset class. We minimize what we collect, we guard what we hold, and we treat privacy as an act of respect — not a compliance checkbox performed for auditors. There is no version of this company that trades on the intimacy of what people tell us.

We believe integrity is the moat.
Health tech is drowning in overclaim — AI that "understands" what it doesn't, outcomes announced before they're measured, decks that promise what the product can't do. We think that's not just wrong; it's bad strategy. Trust is the scarcest resource in this industry, and it compounds for whoever refuses to spend it cheaply. In a market full of hype, the ones who told the truth is the only durable position — and the only one worth having.

We believe this must be built by people who are in it.
Not at arm's length. Not as a market opportunity someone spotted on a spreadsheet. The conviction under this company comes from proximity to the thing itself — from knowing what the 2am actually feels like, and what it means when a tool treats you like a whole person instead of a case. That closeness is not a liability we manage. It's the reason the product tells the truth.


What we refuse

  • We refuse to ship a demo dressed as a product. If it's stubbed, gated, or unproven, it says so — or it stays hidden until it's real.
  • We refuse to gamify despair. No streaks for surviving. No badges for a bad brain day. No false cheer.
  • We refuse to call it "AI-powered insight" when it's a heuristic, or "clinically validated" when it's a hope.
  • We refuse to turn measurement into a leash — no scoring people, no ranking patients, no algorithmic verdicts handed down over clinical judgment.
  • We refuse to mine, sell, or "leverage" the private weight of someone's worst season. Ever.
  • We refuse to bury uncertainty. The error bars are part of the truth.
  • We refuse to grow by any means that requires someone to be reduced.

The flag

We are not the loudest company in behavioral health, and we don't intend to become it. We intend to be the one whose word holds.

So this is the flag, planted plainly: every claim evidenced, every gap admitted, every person met with dignity, every clinician armed with truth instead of theater. If we can't build it honestly, we won't build it. If we can't say it truthfully, we won't say it.

The people we serve have been let down enough — by systems, by stigma, by software that promised understanding and delivered a dashboard. They don't need another miracle.

They need someone who tells the truth and stays.

That's us. That's the company. Hold us to it.


— MoodTrace Health Inc.