A strategy you can back up.
Turning Research Into Strategy is a Claude Code skill for the pile of calls, notes and customer feedback nobody has turned into direction. Your agent lists what you have and how far to trust each piece, counts what the evidence really supports, opens the product and tries it, and only then writes the strategy. It gives up the easy assertion: every number in the finished document can be traced back to the person who said it.
Install it
› /plugin marketplace add hgambas/think-like-a-pm
Then /plugin install think-like-a-pm, which
brings this skill and Context is Queen together. It starts on its own when
you have research that hasn't become direction, or you can call it by name.
The problem
Research goes stale in the gap between hearing it and using it. The calls get recorded, the painpoints get logged, and then a plan gets written from memory. By the time anyone checks, the plan is already ordering the work.
- Twenty calls and no direction. All transcribed, all sitting in a folder. The owner usually calls it a mess.
- “Eleven customers asked for this.” Nobody wrote down what counted as asking, so nobody can check the number.
- “The product already does that.” Nobody has opened it in weeks, and the button everyone remembers has nothing behind it.
How it works
The work runs in six steps, and each one writes its own document: a real file you can open, not an answer in the chat. Nothing can start before the documents it depends on exist. If one is missing, your agent says which, offers to run it first, and asks before starting anything big. Order governs starting, not finishing: any step can be re-run when new evidence lands, and doubling back is normal.
- 1.Take the material in, and grade it
Transcripts, auto-notes, the painpoint tracker, the backlog, old decks, and access to the product. Each source gets a confidence grade as it arrives.
- 2.Unpack each call on its own
One call at a time, sorted into six buckets, with the difference between what they asked for and what they are stuck on written down.
- 3.Inventory everything flat
Every item, in the words it was said in, next to the person who said it. Nothing gets grouped, ranked or interpreted at this stage.
- 4.Check the counts
Write the counting rule, agree it with you, then count. Publish what was claimed next to what the count found.
- 5.Check what's real
Open the product and try each thing it appears to do. Audit every data source twice, against official documentation and a live test.
- 6.Write the strategy, then the backlog
What you can honestly promise to whom, at most two goals, and rules that decide what gets built. Then the ordered work list under those rules.
Before the first step, you and your agent agree in one sentence what the finished work has to let you do. A real one: “Done when the people building leave the working session knowing what they are prioritizing and why.” Anything that doesn't serve that sentence is scope, not progress.
Where you come in
Some decisions are yours and stay yours: the counting rule, the choice between candidate goals, and permission to start a long pass. At each one your agent stops, asks a short question with a few concrete options, and puts its recommendation first. If nobody is there to answer, in a scheduled or background run, it takes its own recommendation, records it as provisional, and raises it the next time you're around. It never settles one of your decisions quietly.
Choose a command
You don't need the command names. Plain requests work and your agent routes them. The commands are for when you want to steer directly:
| You're thinking… | Run |
|---|---|
| “This plan doesn't smell right” | challenge |
| “I have a pile of calls” | discovery |
| “Make sense of this one call” | debrief |
| “Lay out everything we heard” | evidence |
| “Is that number real?” | validate |
| “Does it really do that?” | teardown |
| “Can we even get that data?” | feasibility |
| “What do we build, and why?” | strategy |
| “What's next on the list?” | prioritize |
| “What are we waiting on?” | decision-log |
| “Does this clash with my setup?” | checkup |
Useful pairs: challenge → validate (name the doubt, then settle the count) · strategy → prioritize (direction, then the work it implies).
The idea in one paragraph
Two habits do most of the work here. The first is writing the counting rule before you count, so the count can't drift toward the answer someone was hoping for. The second is opening the product and clicking the thing before anyone promises it, because a button is a claim rather than a feature. Everything else exists to keep those two honest: raw evidence kept raw until it's counted, sources graded by where they came from, and a running record of what was decided, what is open, and who you're waiting on.
Where it has been proven
This method comes from business-to-business work with a single operator-buyer and a small pipeline of named customers. It's untested on consumer products, large pipelines, and markets with no named customers to follow. It also needs research to exist already: it mines evidence, it doesn't generate it. Skip it when direction is settled and you're only sequencing work inside it, or when you want to spar on one strategic question rather than produce the document.
challenge
Point it at a plan you distrust. It names the bet the plan is making, tests that bet against the evidence, and asks for receipts on every claim doing real work.
When to use it
When a strategy, deck or roadmap already exists and is about to drive commitments, or when something in it doesn't sit right. This is often where the whole thing starts. It can run before any other step, and what it finds usually becomes the work list for the rest.
How it works
- 1.It names the core bet
One sentence on the main thing the plan assumes, for example “retention is the wedge”. It confirms that with you first, because everything after aims at whatever gets named here.
- 2.It tests the bet
What would have to be true in the evidence for the bet to hold? Where counted evidence exists, it checks against it. Where it doesn't, it lists the claims that need counting.
- 3.It asks for receipts, claim by claim
Each load-bearing number, quote or capability claim traces back to a named source, or gets marked untraceable where it stands. Nothing is deleted quietly: you can still see what was claimed and what couldn't be backed.
- 4.It compares the live documents
The deck against the strategy, the strategy against your own landing page, all of them against what people said on calls. Every contradiction becomes a numbered open question. Two documents disagreeing is the finding, so it gets written down plainly.
- 5.It closes with a verdict and a work list
Whether the bet survived, which claims held, which didn't, which are still open, and which commands would settle each one.
Try it
That question is where this method came from: a comment on a draft.
A finished challenge reads like this: the draft leads with conversion, but the calls, strictly counted, support retention, with 2 qualifying calls rather than the claimed 7. Two claims trace to no source. The deck and the landing page state different prices.
Pitfalls
- Expecting it to rewrite the plan. It names what doesn't hold. The replacement comes later, once the evidence work is done.
- Letting it half-audit the numbers here. Tracing a claim is this command's job. Strict recounting under a written rule belongs to validate.
- Pointing it at your own decisions. It targets the document's claims and bets. Decisions you made get revisited by you, through the running record.
Collect the research
Three commands get everything into the room and written down: what you have and how far to trust it, what each call actually said, and a flat list of every item next to the person who said it.
When to use it
At the very start, and again whenever new material arrives: a fresh batch of calls, an updated deck. Don't wait for the pile to be complete. Your agent takes what exists and files the gaps as open questions with a name against each one.
discovery: grade every source at the door
It asks for the material one item at a time: call transcripts, auto-notes from tools like Granola or Otter, the painpoint tracker, the backlog, any existing strategy decks, and access to the product itself. If there's no prototype, it asks for whatever is closest to real, like the demo or the spreadsheet the founder actually runs things from.
Every source gets a confidence grade as it arrives:
- a transcript you recorded yourself is the strongest, and can be quoted word for word
- auto-notes with a machine transcript need the wording checked before you quote them anywhere public
- notes with no transcript are the weakest, and can only be paraphrased, because no exact wording exists
The grade is about where a source came from, not how tidy it looks. A beautifully formatted set of notes with no recording behind it is still the weakest grade. The mix gets stated openly, for example “14 of 20 calls are machine-transcribed, 1 is notes-only”. Decks come in as claims to be audited, never as truths to inherit.
debrief: one call, six buckets
Run it right after a call that produced commitments, decisions or a problem statement, or later on any recorded call nobody processed. One call per run. If there's no transcript or notes, it says so and stops, rather than reconstructing the call from memory.
- explicit commitments: who promised what, by when
- things you're blocked on that the other person holds
- their problem, stated precisely, in their own words where the grade allows
- decisions actually reached, dated and attributed
- open questions raised but not resolved
- work the conversation assumes but nobody claimed, with every guess marked as a guess
Then it separates what they asked for from what they're stuck on. Those two often differ, and the difference is the finding. The stated ask gets treated as a hypothesis rather than a work order.
evidence: the flat list
Every painpoint, quote and request from every source, recorded in the original words where the grade allows, with its source, its grade, and an ID so later documents can cite it. One tag per item, and only one: is this a problem in the customer's own world, or friction with your product? Those two get used differently later.
Nothing is merged, themed or ranked here. If your agent catches itself writing “several customers said”, it stops, because that sentence belongs to the counting step. Weak sources stay in the list, marked weak, rather than disappearing. The point is that anyone can re-run the analysis later under a different rule, because nothing was pre-digested.
Try it
Then: “Lay out everything we heard, flat.”
Pitfalls
- Waiting for the full pile before starting. Missing material becomes an open question with a name against it, not a reason to stall.
- Treating the deck as evidence. A deck is a set of claims to check. If you're about to quote one as fact, you wanted validate.
- Grouping items as you go. The moment the list is sorted by theme, every later count inherits those groupings.
validate
Write the counting rule before counting anything. Then publish what was claimed next to what the count actually found.
When to use it
Whenever a number, a quote, or a “the research shows” claim is about to drive a decision: a build order, a pricing move, a promise to a partner. Also any time two live documents disagree. It counts out of the flat list, so if that doesn't exist yet, it offers to build it first.
How it works
- 1.The rule comes first, and you confirm it
A real one: “a call counts for a goal only if the person described the pain in their own world. Market commentary, advice, and descriptions of other people's problems do not count.” Where to draw that line is your judgment call, so it waits for your yes.
- 2.Then it counts, citing item IDs
Borderline cases are listed separately with the reasoning for each, so the whole count can be re-run under a looser rule without redoing the work.
- 3.A second count, for customer fit
Each source is graded against who the product is actually for: fits, marginal, outside, or not a buyer, with a reason. This stops loud non-buyers from outweighing your real segment.
- 4.Claimed next to audited, with a verdict
Then a short paragraph per mismatch explaining which assumption produced the inflated number. That explanation is the useful part, because it stops the same inflation happening again.
- 5.Counter-evidence gets recorded too
If a framing actively failed in three calls, that goes in the table alongside the support.
- 6.Whatever can't be settled goes to the record
As a numbered open question, or as a named person you're now waiting on.
Try it
It will write the counting rule and ask you to confirm it before counting.
Claimed: 11 calls supported matchmaking. Audited under the own-world rule: 3 qualify. A real signal, and a long way from the majority theme it was presented as.
Pitfalls
- Counting first and writing the rule after. The order exists so the count can't drift toward the answer someone wants.
- Deleting the claims that failed. Downgrade them where they stand and show the gap. A quietly deleted claim teaches nobody anything.
- Stopping at “the number was wrong”. Without an explanation of how the wrong number got made, the audit is half done.
Check what's real
Two commands that run before anything is promised: open the product and try it, and audit every data source you would be depending on.
When to use it
Before any plan, promise or backlog item references something the product does or some data you expect to get. Also any time someone says “the product does X” and nobody in the room has recently watched it do X.
teardown: try it, don't read about it
Your agent walks the product screen by screen and records what it saw, not what was intended. Everything the interface implies gets tried. The result is two lists with a hard line between them: what works, and what is a mockup, meaning it looks real but isn't wired up, because it's visual only, hardcoded, or running on fake data. There's no middle column for things that probably work.
Things that work but work badly go on a third, separate list, because fixing something and making something real are different jobs. It closes by saying what it deliberately didn't test and why. An untested capability listed as untested is honest. One left off the list becomes an assumed capability by next week.
It also notes anything the product already holds and does nothing with. In one real run, an event page held 86 registrations and only 10 of those people were members. That meant 76 prospects already sitting in the product, unused, which became the cheapest growth item on the backlog.
feasibility: what each data source will actually give you
The same three questions of every source you might depend on: what can be read, and in how much detail; who grants that access, whether that's the customer, their plan tier, or an approval process at the platform; and what it costs in money, rate limits, review queues or licence terms.
Answers come from official developer documentation, the live product, or a real call to the interface. Never from memory or from a vendor's marketing page, because pricing pages and feature grids are claims rather than evidence. Two passes run over every source, and the second one is adversarial: its job is to overturn the first. Where they disagree, the original documentation settles it. Where nothing settles it, the claim stays marked unverified, with the specific action that would settle it, like “ask the vendor in writing” or “one manual check in the browser”.
Two distinctions save a lot of wasted building. No documentation is not the same as no capability. And a gate on the interface is not a gate on the data: a plan tier that blocks the live connection may still allow an export, which turns “we can't serve them” into “upload it instead”.
Corrections are ranked by how much each one changes a decision, not by topic, so the finding that kills a build comes first. Then it asks you to spot-check the top few yourself. Two agent passes are strong evidence and still not enough: in the work this came from, the human check reversed three assumptions about platforms. It closes by naming which sources change often and must be re-checked before anything depends on them.
The shape of a real correction: the renewal date moved to a different place in the latest version of the interface, so code written to the old shape stops noticing renewals, without any error to warn you.
Try it
Then: “Audit every data source we'd be depending on.”
Pitfalls
- Reading the interface as the capability. A button is a claim, not a feature. Click it.
- Walking only the happy path. The demo route is the best-lit part of the product. Walk the rest.
- Auditing only the sources you like. The audit covers every source a promise might rest on, including the awkward ones.
strategy
The main one. What you can honestly promise to whom, the two goals the work serves, and the rules that decide what gets built. It sits on top of every step above and offers to run whichever ones are missing.
When to use it
When someone asks for “the strategy”, “direction”, or “what do we build and why”. Also to re-derive it after new evidence lands. Before it writes anything it checks for the earlier documents, names the missing ones, and offers to run them in order, asking you before each. It won't write a strategy over a gap, and it won't start a five-step run without your yes.
How it works
- 1.It groups your data sources by the job each one does
Not by vendor category. A source can sit in more than one group. This grouping stays internal, and no customer ever needs to learn it.
- 2.It derives at most two goals from the counted evidence
Smaller goals sit under those two. It presents the candidates to you as a choice with its recommendation first, because goals are your call. The two aren't ranked once and for all: which one to lead with is decided per customer.
- 3.It writes the coverage-to-promise table
The centre of the document. Each row is what a customer has connected, and says what you can honestly answer from it, what you can't, and what that's good enough for.
- 4.It writes the build rules
A procedure for deciding, rather than a roadmap of dates.
- 5.It numbers the open questions and refers to them in the body
So a reader meets each question at the point where it bites, not only in a list at the end. It closes with the evidence base: the named people and their roles, and every source with its date and grade.
The coverage-to-promise table
This is the table that turns a sales conversation into a diagnosis. Every row states what you cannot answer as well as what you can, and at least one row is a refusal:
| What's connected | What you can honestly answer | Good enough for |
|---|---|---|
| Revenue only | Who stopped paying, and who is about to renew. Not why, and not who might join next. | A thin version of retention work. Nothing else. |
| Revenue and chat | Who is going quiet compared to their own normal, before it costs a renewal. | Retention, for communities that live in chat. |
| Chat only | Who went quiet. Not whether they matter, or what to do about it. | Nothing on its own. Don't onboard anyone on this alone. |
Each promise carries the recipe behind it. And there's a hard rule for saying no: if the outcome a customer wants has never been recorded in their data, the product would be guessing. So the answer is to decline that goal and tell them what's missing.
The build rules
A procedure, not a roadmap. The reference shape runs like this. When a customer says yes, run the coverage check, and the gap in their coverage names the integration to build. Check the data-source audit: can that source really yield what's needed, and behind which gate? Across the whole confirmed pipeline, build whatever frees the most blocked customers. And if no confirmed customer is blocked on something missing, build nothing new and onboard on what already exists. That last rule is the one most roadmaps can't bring themselves to say.
How the documents read
Everything this skill writes is aimed at a smart reader outside your team who wasn't in the room. Everyday words, complete sentences, and a plain explanation of any internal term the first time it appears. Numbers say what they mean, and people get introduced.
BeforeMRR $5K (committed). Rev evidence: 0 priced. Deck targets n/a, learning goal only.
AfterThe deck's $5,000 monthly revenue is committed, not collected. No customer has paid a price yet. We inherit the deck's learning goal, and not its revenue targets.
At the end, the draft goes through three passes. Anything defending the document rather than informing the reader is cut. Every surviving sentence has to survive the question “how do you know?”. Then the order and headings are made simple enough for a stranger to navigate. The finished document should be shorter than the draft it came from.
Try it
It will list any missing steps first and ask before running them.
Pitfalls
- Writing strategy over a missing step. If the counts were never checked, the strategy inherits the inflation.
- A third goal arriving quietly. New goals enter as a recorded decision, never through a promising feature idea.
- Promising from the deck. A deck claim that hasn't been counted doesn't get promised. You can inherit a deck's learning goals, but never its targets.
prioritize
Order the work so that every item earns its place and names the goal it serves. This is where a strategy either becomes instructions or quietly falls apart.
When to use it
Once the strategy exists, and again to re-order when new customers confirm or new evidence lands. The list is generated from the strategy and the product teardown: mockups to make real, rough edges to fix, and gaps between the product and the goals. An existing backlog gets audited against the rule and merged where it survives, never simply inherited.
How it works
- 1.The qualifying rule goes at the top, and you confirm it
The reference shape: “a capability counts only if it helps the customer act on the goal. A chart nobody acts on does not qualify.” One real customer had already rejected the softer version: “I don't need a platform to tell me that people are silent.”
- 2.Every item is tagged
With the goal it serves, and with its kind: make real (a mockup becoming true), net new, or fix.
- 3.The order is stated, and so is the logic behind it
One line at the top. Blocking issues first, then the cheapest proof of a goal with thin evidence, is a reasonable default. Whatever the logic is, it gets written down, so the next re-order isn't a fresh argument.
- 4.What's deliberately excluded, and why
An exclusion with a reason can be argued again later, while a silent omission can't be. A real one: introductions were kept off a backlog “because building them now would add a third goal before we have decided to take one on.”
- 5.Dependencies are checked against the data-source audit
Any item resting on an unverified data claim gets that claim's settling action scheduled before or alongside it.
Try it
It will propose the qualifying rule and wait for you to agree it.
A good test of an item: “dashboard showing customer activity” usually fails the rule, while “rank these people, with a drafted action for each” usually passes.
Pitfalls
- Items that describe analysis instead of capability. If nobody can act on the output, it doesn't qualify.
- An order with no stated logic. Unwritten ordering rules get re-argued every time the list changes.
- Sneaking in a third goal. An item serving no listed goal is either cut, or logged as a proposal for a new goal.
The running record
One document that always answers three questions: what have we decided, what is genuinely open, and who are we waiting on. Every command adds to it as it works, and it's the first thing to read when you pick the project back up.
When to use it
Call decision-log directly to review where things stand, record a decision made outside a command, close a question that got settled, or tidy up after a messy week.
What's in it
Three separate registers. The split is by who can close each item, which is what makes the document usable:
- Decisions locked, each dated and attributed to the person who made it. A locked decision is only reopened by that person.
- Open questions, numbered so other documents can point at them, each stating what would settle it.
- Waiting on a named person, with the item, the name, and the date the wait started. A wait long enough to block work gets raised with you rather than aging quietly.
A few housekeeping habits keep it honest. Terms are defined once at the top and used consistently, including what counts as a confirmed customer, because most contradictions between documents turn out to be two definitions wearing one word. “By Friday” becomes a real date. A settled question moves into the decisions list with its date rather than being deleted, because the trail is the point. Corrections are recorded plainly, including the ones that reverse your own earlier framing.
checkup: a one-time scan of your setup
Run it once after installing, and again after you add other research or strategy skills. It looks at the other skills and standing instructions on your machine and reports four things: skills that would fire on the same kind of request, instructions that contradict this method, files big enough to eat context every session, and whether your total skill list has outgrown the space the tool gives it. Each finding comes with a suggested fix.
It's report-only and local-only. It never edits, disables or deletes anything, nothing it reads leaves your machine, and any location it couldn't read is listed as unchecked rather than guessed at. In a chat app with no files to scan, it says so and reviews what it can see instead.
Try it
Or, right after installing: “Scan for anything that clashes with this skill.”
Pitfalls
- Registers blurring together. “Waiting on Jules” filed under open questions makes it nobody's job.
- Decisions with no owner or date. An unattributed decision gets re-argued, and an undated wait never escalates.
- Using it as a to-do list. Work items belong in the backlog. This tracks what you know, not what you're doing.