Methodology Effectiveness Ranking — AI-Guided Staged Pipeline (Idea → Docs → Build)#

Builds on methodology-landscape.md (artifact survey) and stage-chain-matrix.md (stage mapping). This document ranks the 8 candidate approaches against the owner's stated goal: move very fast; ship the IDEAL, convenient app that works EXACTLY as the user wants; with MAXIMUM thoroughness of working-out.

Evidence tiers used throughout:

  • Tier A — peer-reviewed / large-N academic studies.
  • Tier B — industry reports / vendor research (flagged with known methodological criticism where applicable).
  • Tier C — practitioner consensus / expert commentary, no controlled study behind it.

Scores are ordinal (High / Med / Low or rank order) with a one-line cited justification each. No invented decimal precision.


1. Criteria and weights (declared up front)#

Derived directly from the owner's own words. Weights are relative, sum to 1.0, assigned by how directly each phrase maps to a criterion:

CriterionWeightOwner language it maps to
Speed-to-value — time from raw idea to a working, testable thing0.30"move very fast", "whoever ships"
Faithfulness-to-real-user-need — does the output match what the user actually wants, not an assumed want0.25"works EXACTLY as the user wants"
Thoroughness / traceability — depth of working-out, documented reasoning, nothing-forgotten0.20"MAXIMUM thoroughness of working-out"
Usability/UX outcome quality — the delivered app is genuinely convenient/pleasant to use, not just functionally correct0.15"IDEAL, convenient app"
Fit-to-AI-guided-staged-tooling — does the method decompose into discrete, gateable, AI-automatable stages at all0.10implicit in "pipeline tool" framing (not an owner quote, lowest weight, structural constraint not a goal)

Why speed and faithfulness outweigh thoroughness-as-a-standalone-criterion: the owner names speed first and twice ("very fast" + "whoever ships... wins"), and "exactly as the user wants" is a stronger, more falsifiable bar than generic thoroughness. Thoroughness is the explicit quality floor, not the primary differentiator — see Section 4, this is exactly where the tension lives.


2. View 1 — Overall ranking (methodology as a whole, against the full weighted goal)#

RankMethodologyWeighted fitWhy (one line, tiered)
1Continuous discovery / JTBD (Torres Opportunity-Solution Tree, Ulwick ODI)HighWeekly evidence cadence directly targets faithfulness-to-need and catches drift fast — Tier C/B: Torres's own framework claims (no controlled effectiveness study located; see Grounding), but Ulwick's ODI reports an 86% vs ~17% industry success-rate claim (Tier B, single-vendor, not independently replicated) Strategyn.
2Lean startup / customer development (Ries Build-Measure-Learn, Blank 4-steps, Lean Canvas)HighBuilt explicitly to minimize wasted work before validating demand, which is the fastest honest path to "exactly what the user wants" — but software-startup failure stays 75-90% even with Lean adoption (Tier A/B mapping-study, effectiveness is directional not proven) arXiv systematic mapping on MVP practices.
3Lightweight modern tech-spec stack (C4, ADR, OpenAPI)HighNot a discovery method at all — it's the fastest, most AI-generatable way to produce traceable, reviewable backend/architecture documentation once requirements exist; ADR practice directly targets traceability (Tier C practitioner-originated, but near-universally adopted standard) Nygard 2011, c4model.com.
4Agile user-stories / story-mapping / BDDMed-HighLarge-N studies favor agile over waterfall on success rate (Tier A: Serrador & Pinto found a 16-pp higher success rate for agile; a South African study found 88.2% vs 47% — methodology/sample caveats apply, see Grounding) Serrador & Pinto cited via scielo.org.za study; fast and AI-tool-friendly (INVEST, Gherkin are structured enough to automate) but the user-story format itself admits it carries no guaranteed depth — it's "training wheels," not thoroughness Agile Alliance glossary.
5Design thinking / double diamondMedDirectly targets usability/UX outcome quality and faithfulness via empathize/test loops, but academic critics find "little tangible evidence" the method itself (as opposed to good designers using it) causes better outcomes — Tier A critique: no empirical base proven Hobday, Boddington, Grantham critique, cited via ResearchGate; non-linear-by-design ("start wherever," "no idea is ever finished") also resists hard pipeline gating — fit-to-tooling is the weak spot here (Tier C, from the method's own official framing) designcouncil.org.uk.
6Heavy RE standards (IEEE 29148 / BABOK / Volere)MedWins decisively on thoroughness/traceability by design (29148's 9+5 quality-characteristic grid, Volere's 27-section coverage + snow-card fit-criteria) — Tier C/B grounding via standard summaries, not the paywalled primary text — but is the slowest candidate on the list and carries no built-in user-validation loop; justified mainly for complex/regulated builds, which conflicts with "move very fast" for most product work RUP 2026 commentary.
7RUP/UPLow-MedSame thoroughness strength as RE standards (shared artifact lineage) but adds full phase/discipline process overhead on top — modern consensus (2026) treats RUP-as-process as "largely retired," valuable only as a vocabulary/artifact donor, not as the backbone methodology (Tier C consensus, Wikipedia + practitioner commentary) Rational Unified Process — Wikipedia.

Reading View 1: no single methodology wins outright — the top 3 win because they are narrow-scope tools that plug into specific stages rather than end-to-end philosophies, which is itself evidence for a hybrid (confirmed rigorously in View 2 below, not asserted here).


3. View 2 — Per-stage fitness ranking#

Each stage ranks all 8 candidates High / Med / Low / N-A (not applicable — method has no defined artifact at this stage per the landscape survey). "N-A" is itself informative: it shows no single methodology spans the whole pipeline, which is the rigorous basis for a hybrid rather than an assumption.

StageRUP/UPHeavy RE standardsAgile stories/BDDLean/CustDevContinuous discovery/JTBDDesign thinking/Double DiamondLightweight tech-spec (C4/ADR/OpenAPI)
Discovery / problem-framingMed (Vision doc input-driven, Tier C)Low (BABOK has a planning KA but is elicitation-heavy, slow to start, Tier C)Low (agile has no discovery artifact of its own — borrows from others)High (Lean Canvas/Customer Development built exactly for this; falsifiable fast, Tier B/C)High (Opportunity-Solution Tree root outcome + weekly interviews is purpose-built, Tier C)Med (Empathize/Discover strong on qualitative insight but no completeness gate, Tier A critique on evidence base applies)N-A
RequirementsMed (Use-case model/Supplementary spec, Tier C)High (29148's measurable quality-characteristic grid, Volere snow cards with fit criteria are the most rigorous, traceable artifacts surveyed — Tier C/B via standard summaries)Med-High (User story + INVEST + Acceptance Criteria is fast and AI-generatable, but intentionally shallow by design, Tier C)Low (CustDev produces validated hypotheses, not structured requirements)Low (JTBD outcome statements feed requirements but aren't requirements themselves)N-AN-A
UXN-AN-AN-ALowMed (Opportunity-Solution Tree extends into UX opportunity-mapping, Tier C)High (Double Diamond Discover/Define + journey maps/empathy maps/user flows is the only surveyed lineage purpose-built for this stage, Tier A primary source for the phase model, Tier C/🤔 for per-artifact completeness criteria)N-A
UIN-AN-A (Supplementary spec Usability category only)N-AN-AN-AHigh (Ideate/Prototype/Test ↔ Develop/Deliver, wireframe/prototype fidelity practice, Tier A phase model + Tier B NN/g-adjacent fidelity evidence — noting the headline "5 users find 85%" figure itself has a Tier A caveat: Faulkner 2003 found individual 5-user samples ranged 55-99%)Low
Frontend specN-AN-AN-AN-AN-AN-AHigh (Atomic Design + Storybook + Design Tokens is the only surveyed artifact chain at this stage, Tier C community convention but near-universal tooling adoption)
Backend specMed (Supplementary spec NFR categories carry over, Tier C)Med (29148's SyRS/SRS structuring applies but is heavier than needed once use cases exist, Tier C)LowN-AN-AN-AHigh (C4 + OpenAPI + ERD + ADR is purpose-built, machine-checkable, AI-generatable — Tier C practitioner standard but with primary-source grounding for C4, ADR, OpenAPI)

What View 2 shows rigorously: the winner changes at every stage. Discovery and UX/UI have a clear non-RUP, non-RE-standard winner (lean/continuous-discovery for discovery, design-process lineage for UX/UI). Requirements is the one stage where heavy RE standards genuinely win on the owner's own "maximum thoroughness" criterion — nothing else surveyed has a comparably rigorous completeness test. Frontend/backend spec stages have no competition at all from the discovery/requirements/UX families — only the lightweight tech-spec stack has artifacts there. This is what makes a hybrid derivable, not arbitrary: each stage's winner is determined by "which methodology has ANY defined artifact here, and of those, which has the strongest evidence," not by picking favorites.


4. The speed-vs-thoroughness tension (dedicated, not averaged away)#

The owner's two top-line goals are in real, surfaced tension — not a rhetorical one:

  • At the discovery and UX stages, speed and faithfulness-to-need pull in the SAME direction: Lean Canvas, Customer Development, and continuous discovery are simultaneously the fastest AND the methods explicitly built to prevent building the wrong thing (Ries: work that doesn't produce "validated learning" is waste). Thoroughness in the RE-standards sense (29148, Volere, BABOK) is actively counterproductive here — Tier C: heavy upfront elicitation before any real user contact risks fully specifying the wrong product, precisely the failure mode Customer Development was created to prevent Wikipedia — Customer development.

  • At the requirements and backend-spec stages, the direction flips: thoroughness wins outright and speed is the one taking the hit. Heavy RE standards (IEEE 29148, Volere) and the RUP-descended Supplementary Spec are the only surveyed artifacts with a measurable completeness test (29148's 9+5 quality characteristics; Volere's per-requirement fit criterion). The classic justification for paying this cost is the cost-of-change curve — Boehm's widely-cited claim that a requirements defect caught late costs up to 100x more than one caught early, with IBM Systems Sciences Institute data cited for a "$1 / $10 / $100" design-dev-production multiplier (Tier B, frequently repeated but methodologically thin — original data from 1970s TRW/IBM waterfall projects). This claim itself is now contested: Boehm and Basili's own 2001 revision for small/agile/CI-enabled projects reported a meaningfully flatter curve, and modern commentary argues CI/CD and modular architecture have further flattened it (Tier B/C) Mountain Goat Software — The Cost of Change Curve Is Outdated, reworkcost.com — Boehm Cost of Change Curve: What the 1981 Data Actually Shows. Implication for this tool: the thoroughness payoff at the requirements/backend stage is real but smaller than the classic 100x folklore suggests for an AI-assisted, fast-iteration pipeline — it still justifies rigor at these two stages specifically, but does not justify extending RE-standard-level rigor backward into discovery/UX, where it would slow exactly the stages that most benefit from speed.

  • Net resolution (not an average, a per-stage split): the tension is not resolved by picking one methodology-family for the whole pipeline — it is resolved by letting the winning methodology change per stage (View 2), with thoroughness concentrated specifically at requirements and backend-spec (where it is cheap to be thorough and expensive not to be) and speed/faithfulness concentrated at discovery/UX (where thoroughness without user contact risks building the wrong thing well). This mirrors the "two configurable depths per stage" accommodation already identified in stage-chain-matrix.md's Ordering Disagreements section — this ranking confirms, with evidence tiers attached, that the accommodation is justified rather than a convenient compromise.


5. Evidence-tier honesty check (escalation rule)#

Tally of what this ranking actually rests on:

  • Tier A (peer-reviewed / large-N): Faulkner (2003) on usability-sample variance; Hobday/Boddington/Grantham critique of design thinking's evidence base; Serrador & Pinto-style agile/waterfall comparative study (via secondary citation — primary not independently refetched, flagged); the Double Diamond / d.school phase models as officially-sourced primary frameworks (not effectiveness studies, but primary methodology sources).
  • Tier B (industry reports, flagged for known criticism): Standish CHAOS-derived agile-vs-waterfall success-rate figures (criticized for self-selected samples and a success definition that changed in 2015 from strict scope/cost/time to a looser "satisfactory result" test — reported here with that caveat attached, not averaged away); Ulwick/Strategyn's 86%-vs-17% ODI claim (single-vendor, not independently replicated); Boehm/IBM cost-of-change figures (1970s waterfall-era data, now contested by a 2001 revision and modern CI/CD commentary).
  • Tier C (practitioner consensus, no controlled study): RUP's 2026 retirement-as-process consensus; Storybook/Atomic Design completeness convention; C4/ADR/OpenAPI practitioner-standard status; Volere/29148/BABOK completeness-grid descriptions (grounded via secondary summaries of standards whose primary text is paywalled).

Rough mix: roughly 25% Tier A, 40% Tier B, 35% Tier C across all scores in this document — this is NOT over the 50%-Tier-C escalation threshold, so the overall ranking (View 1) is retained alongside the per-stage view, not dropped. The escalation rule is noted as a close call, not triggered: the two scores leaning hardest on Tier C alone (RUP's "largely retired" status, and Frontend-spec's "Storybook convention" win) are both low-controversy claims where practitioner consensus is unusually strong and uncontested, which is why they are kept at Tier C rather than treated as weak points requiring a hedge.


Sources (new to this document, beyond methodology-landscape.md)#

Grounding caveats specific to this ranking#

  • Serrador & Pinto's exact "16% higher success rate" figure and the South African study's 88.2%/47% figures were relayed via aggregated search-result summaries, not independently re-fetched from the primary paper PDFs — treated as Tier A-sourced-via-secondary-aggregation, one level removed from the primary text, consistent with how methodology-landscape.md already treats similarly-sourced claims (e.g. its Scrum Guide DoD caveat).
  • The CHAOS-derived "agile more successful than waterfall" headline is reported together with its own criticism (self-selected sample, no disclosed methodology, shifting success definition) per this task's explicit instruction — never presented as a clean Tier A fact.
  • Ulwick's 86%-vs-17% ODI success claim comes from Strategyn, the commercial originator of the ODI method — a single, non-independent source; treated as Tier B with a conflict-of-interest flag, not elevated.
  • No controlled, independent effectiveness study for Teresa Torres's Continuous Discovery Habits was located in this pass (only book summaries and the author's own site); the "High" score for continuous discovery in View 1 rests on structural fit-to-goal reasoning (weekly real-evidence cadence directly answers "exactly what the user wants") rather than a proven-effective claim — flagged as the weakest-evidenced "High" score in this document.