AI Output Clinic · Overview

Find where it broke.

Generic copy is what writing sounds like when it treats the reader as a target. AI Output Clinic is a Claude Code skill for the moment your AI-written content comes back that way: flat, generic or off-tone. It traces the whole chain, from the raw input to the prompt to the screen, and finds the layer that actually caused it. Then it gets your standard out of you using real examples, and turns every failure into a check so the same problem cannot come back.

Coming soon

This skill isn't out yet. It's being used and sharpened on real work first, the same way Context is Queen was.

Follow along to hear when it lands:

In the meantime, Context is Queen is available today.

The problem

When AI content disappoints, the instinct is to rewrite the prompt. That skips the real work. The cause almost always sits in one specific place: the input data the model was given, the prompt itself, the shape the output has to fit, or the screen it lands on. A prompt tweak does nothing about a cause that lives in one of the other three.

  • “It reads so generic.” You rewrite the prompt three times, get three flavours of the same flat paragraph, and stop asking why.
  • “I can't explain it, it's just not right.” You get asked what tone you want. No adjective you can think of makes the next draft better.
  • “It was fine yesterday.” A prompt change fixed one thing and quietly broke another, and nothing was watching.

How it works

  1. 1.
    Diagnose the chain before rewriting anything

    Four layers can cause bad output. Your agent reads all four and names the one at fault, with evidence.

  2. 2.
    Get the spec out of you with examples

    You see the real input as a numbered list, then three or four full drafts. Your pick becomes the written rule.

  3. 3.
    Take one field all the way, then generalise

    Fix a single piece of output properly. Then work out which of its rules apply to everything else.

  4. 4.
    Mock the target inside the real screen

    Hand-write the version you wish you had, and see it rendered with rich data, moderate data and none.

  5. 5.
    Describe who the agent is, if it's a voice

    For personality work, a description of the agent beats a rule list. Every extra “never” takes something away.

  6. 6.
    Turn every failure into a check

    Models vary run to run. Each failure becomes an automatic test that ships with the fix.

Where to start

There are no commands to learn. Say what's wrong in your own words and the skill routes it. These links exist for when you already know which part you need:

You're thinking…Go to
“This is bad and I don't know why”Diagnose the chain
“I'll know it when I see it”Get the spec out of you
“Everything on this page needs fixing”Take one field all the way
“I want something to hold it against”Mock the target
“Our agent sounds like a policy document”Voice and personality
“It was fine yesterday”Pin every failure

Useful pairs: diagnose → spec (know the cause, then say what good looks like) · mock → pin (build the bar, then make it permanent).

How a session runs

The whole thing is built for someone who reacts better to examples than to abstract questions. So it works this way:

  • one thing at a time, never a batch
  • the plan first, then your go, then the work
  • mockups you can react to instead of questions about what you want
  • numbered options, so you can point at them out loud
  • plain English back, whole answer first and short

The idea in one paragraph

Bad AI output is a symptom, and symptoms are misleading. “Too dry” can be caused by a one-sentence cap buried in a prompt, or by a set of output fields that chops a story into boxes, or by a screen that truncates. Rewriting the prompt is the one move that feels productive and often changes nothing. This skill spends its first minutes finding where, gets a real standard out of you with drafts rather than adjectives, and then pins the fix in place with tests so you only have to win the argument once.

AI Output Clinic · Find the cause

Diagnose the chain

Before a word of the prompt changes, find the layer that caused the problem. There are four candidates: the raw input data, the prompt instructions, the shape the output has to fit, and how the screen renders it.

When to use it

The moment output disappoints. “Too dry”, “too generic”, “not rich enough” are all symptoms. Each one has a cause, and the cause is often not where you would look first.

How it works

Your agent walks the chain end to end and reports which layer is at fault, with the evidence that says so. Three causes come up again and again:

  • Hidden length caps. A “one sentence” instruction buried in a formatting block, or in the description of a single output field. It flattens everything and nobody sees it.
  • Verbatim copying. Output that lifts one summary line straight from the input instead of drawing across the whole of it.
  • Structural causes. When the output is chopped into single-idea boxes, the model has no choice but to write stacked one-liners. No prompt tweak fixes that.

There is one standing rule. If you raise the same hunch twice, even after your agent dismissed it the first time, it goes and proves it wrong with evidence from every layer. That rule earned its place: a real “one sentence” cap turned up exactly this way, after the screen had already been cleared of suspicion by mistake.

Try it

“Every brief it writes reads the same. Find out why before changing the prompt.”
Or: /ai-output-clinic

A good diagnosis names one layer and shows you the line that proves it. If the answer is “let's try a different prompt and see”, the chain hasn't been read yet.

Pitfalls

  • Rewriting the prompt first. Three of the four layers cannot be fixed by a prompt, however well written.
  • Clearing a layer on a hunch. Ruling something out without evidence is how a real cap stays hidden for weeks.
AI Output Clinic · Set the bar

Get the spec out of you

You know good output when you see it, and you can't always write the rule down. This turns what you react to into a written standard the model can be held to.

When to use it

After the diagnosis, when you know which layer to change and now have to say what good actually looks like. Also any time you catch yourself reaching for words like “warmer” or “richer”.

How it works

  1. 1.
    You see the real input

    Your agent shows you the actual raw material the model had, as a numbered list. You can point at it out loud: “include 3, 7 and 11.”

  2. 2.
    It drafts three or four full versions

    Not fragments. Complete versions of the real output, each built around a different lead idea.

  3. 3.
    Your pick becomes the template

    Whichever you choose, plus whatever you say about why, gets written down as the rule the model follows.

One example beats ten adjectives. If you can hand over roughly the paragraph you wish it had written, typed badly and half finished, that is the fastest input there is.

Try it

“Show me the research it was working from, then draft me four versions of this summary.”
Each version led by a different idea, so you have something to choose between.

Adjectives don't survive the trip into a prompt. A version you picked, and said one sentence about, does.

Pitfalls

  • Being asked “what tone do you want?” That question has no good answer. Ask for drafts to react to instead.
  • Reviewing fragments. Half a sentence tells you nothing about how the whole thing reads.
AI Output Clinic · Set the bar

Take one field all the way

Fix a single piece of output completely first. Then work out which of its rules belong to everything else the AI writes.

When to use it

When more than one thing on the screen is written by AI and you're tempted to fix them all at once.

How it works

One field goes the whole way through: diagnosed, spec'd, fixed, checked. Only then does your agent pull out the rules that might travel. The ones that usually do:

  • no hidden length caps, so length follows the evidence available
  • draw on the whole input, never lift a line from it
  • weave the facts around one lead idea instead of stacking sentences
  • on thin data, say less rather than pad

Then it looks at every other AI-written thing in the product and answers one question out loud: one shared fix to the prompt, or field by field? Some rules travel in a single edit. Some fields need their own pair of examples, one good and one weak, and no shared rule will do it for them.

Try it

“We fixed the summary. Which of those rules apply to the other three things on this page?”
The answer is a real fork: one shared edit, or field by field.

One field done properly teaches you more than four half-fixed, and it's the only way to tell which change did the work.

Pitfalls

  • Fixing everything at once. Nothing gets proven, and you can't tell afterwards which edit helped.
  • Assuming one prompt edit covers it. Some fields only improve with their own examples.
AI Output Clinic · Set the bar

Mock the target

Write the output you wish you had, by hand, and see it inside a faithful copy of the real screen. That mock becomes the bar the model is held to.

When to use it

Once you know what good looks like and want something concrete to measure against. Also whenever the screen has to survive thin data.

How it works

The ideal version is written by hand, using only real facts from the real input. Nothing invented. It gets rendered inside a mock of the actual screen at three levels of data richness: rich, moderate and empty. The mock is labelled clearly as hand-written so nobody later mistakes it for something the model produced.

Empty states get the same care rather than being left to the model. The fallback copy is written by hand, honest about what's missing, pointing at something you can do about it, and reviewed string by string.

Try it

“Write the version you wish this card said, and show it to me in the real layout: full data, some data, and none.”
Real facts only, and labelled as hand-written.

If the hand-written version is hard to write, the standard isn't settled yet. Go back and draft variations until it is.

Pitfalls

  • Mocking with invented facts. It sets a bar the real data can never reach, and everyone chases it for weeks.
  • Leaving the empty state to the model. Thin data is exactly where the writing matters most.
AI Output Clinic · Set the bar

Voice and personality

When the output is an agent's voice, describe who that agent is. A list of rules produces something that reads like a compliance memo.

When to use it

When you're setting or changing how an agent sounds, or when a personality you used to like has gone flat and you can't say when it happened.

How it works

The description says who the agent is. Boundaries stay few, plain, and stated once. Two tests decide whether it's working:

  • Read it aloud. Does it sound like telling someone who they are, or like a compliance memo?
  • Count the “never” statements. Prohibitions are a measure of bloat. A personality dies by collecting one more “never” every time something goes wrong.

One more thing to watch for: a voice defined in several independent places. Two separate descriptions of the same agent drift apart on their own, without anyone editing either of them.

Try it

“Our agent sounds flat. Read its personality out loud and count the nevers.”
Also worth asking: is its voice described in more than one place?

A voice you can read aloud and recognise as a person is working. A voice that reads like policy has been patched too many times.

Pitfalls

  • Fixing tone by adding another rule. Each reactive “never” takes a little more of the character away.
  • Keeping the voice in several places. It drifts by standing still.
AI Output Clinic · Make it stick

Pin every failure

Models vary from one run to the next. Every failure you see becomes a fixed, automatic check, so the same problem can't creep back the next time someone edits the prompt.

When to use it

Every time you change a prompt, and every time you catch a bad output. This is the step that makes the rest of the work permanent.

How it works

Two rules come first. Never conclude anything from a single run: models vary even at their steadiest setting, so one good result proves nothing and one bad result proves nothing either. And never chase a failure by running it again and hoping it passes. Instead, each failure you observe becomes a check that runs by itself:

  • a fact attributed to the wrong source
  • a field that came back blank
  • output shorter than it should be
  • output that matches a line of the input word for word

The checks go into whatever testing setup your project already has, and they ship together with the prompt change that fixed them. The fix and its proof travel as one thing.

Try it

“That blank field we just saw. Turn it into a check before we change anything else.”
Then ship the check with the fix, not after it.

Running it again until it passes is not a fix. It's a coin flip you got lucky on, and it will land the other way in front of a customer.

Pitfalls

  • Judging from one run. One good result is not evidence. Neither is one bad one.
  • Shipping a prompt fix with nothing behind it. The next change quietly undoes it and nobody notices for a month.