Methodology Effectiveness Ranking — AI-Guided Staged Pipeline (Idea → Docs → Build)#
Builds on methodology-landscape.md (artifact survey) and stage-chain-matrix.md (stage mapping). This document ranks the 8 candidate approaches against the owner's stated goal: move very fast; ship the IDEAL, convenient app that works EXACTLY as the user wants; with MAXIMUM thoroughness of working-out.
Evidence tiers used throughout:
- Tier A — peer-reviewed / large-N academic studies.
- Tier B — industry reports / vendor research (flagged with known methodological criticism where applicable).
- Tier C — practitioner consensus / expert commentary, no controlled study behind it.
Scores are ordinal (High / Med / Low or rank order) with a one-line cited justification each. No invented decimal precision.
1. Criteria and weights (declared up front)#
Derived directly from the owner's own words. Weights are relative, sum to 1.0, assigned by how directly each phrase maps to a criterion:
| Criterion | Weight | Owner language it maps to |
|---|---|---|
| Speed-to-value — time from raw idea to a working, testable thing | 0.30 | "move very fast", "whoever ships" |
| Faithfulness-to-real-user-need — does the output match what the user actually wants, not an assumed want | 0.25 | "works EXACTLY as the user wants" |
| Thoroughness / traceability — depth of working-out, documented reasoning, nothing-forgotten | 0.20 | "MAXIMUM thoroughness of working-out" |
| Usability/UX outcome quality — the delivered app is genuinely convenient/pleasant to use, not just functionally correct | 0.15 | "IDEAL, convenient app" |
| Fit-to-AI-guided-staged-tooling — does the method decompose into discrete, gateable, AI-automatable stages at all | 0.10 | implicit in "pipeline tool" framing (not an owner quote, lowest weight, structural constraint not a goal) |
Why speed and faithfulness outweigh thoroughness-as-a-standalone-criterion: the owner names speed first and twice ("very fast" + "whoever ships... wins"), and "exactly as the user wants" is a stronger, more falsifiable bar than generic thoroughness. Thoroughness is the explicit quality floor, not the primary differentiator — see Section 4, this is exactly where the tension lives.
2. View 1 — Overall ranking (methodology as a whole, against the full weighted goal)#
| Rank | Methodology | Weighted fit | Why (one line, tiered) |
|---|---|---|---|
| 1 | Continuous discovery / JTBD (Torres Opportunity-Solution Tree, Ulwick ODI) | High | Weekly evidence cadence directly targets faithfulness-to-need and catches drift fast — Tier C/B: Torres's own framework claims (no controlled effectiveness study located; see Grounding), but Ulwick's ODI reports an 86% vs ~17% industry success-rate claim (Tier B, single-vendor, not independently replicated) Strategyn. |
| 2 | Lean startup / customer development (Ries Build-Measure-Learn, Blank 4-steps, Lean Canvas) | High | Built explicitly to minimize wasted work before validating demand, which is the fastest honest path to "exactly what the user wants" — but software-startup failure stays 75-90% even with Lean adoption (Tier A/B mapping-study, effectiveness is directional not proven) arXiv systematic mapping on MVP practices. |
| 3 | Lightweight modern tech-spec stack (C4, ADR, OpenAPI) | High | Not a discovery method at all — it's the fastest, most AI-generatable way to produce traceable, reviewable backend/architecture documentation once requirements exist; ADR practice directly targets traceability (Tier C practitioner-originated, but near-universally adopted standard) Nygard 2011, c4model.com. |
| 4 | Agile user-stories / story-mapping / BDD | Med-High | Large-N studies favor agile over waterfall on success rate (Tier A: Serrador & Pinto found a 16-pp higher success rate for agile; a South African study found 88.2% vs 47% — methodology/sample caveats apply, see Grounding) Serrador & Pinto cited via scielo.org.za study; fast and AI-tool-friendly (INVEST, Gherkin are structured enough to automate) but the user-story format itself admits it carries no guaranteed depth — it's "training wheels," not thoroughness Agile Alliance glossary. |
| 5 | Design thinking / double diamond | Med | Directly targets usability/UX outcome quality and faithfulness via empathize/test loops, but academic critics find "little tangible evidence" the method itself (as opposed to good designers using it) causes better outcomes — Tier A critique: no empirical base proven Hobday, Boddington, Grantham critique, cited via ResearchGate; non-linear-by-design ("start wherever," "no idea is ever finished") also resists hard pipeline gating — fit-to-tooling is the weak spot here (Tier C, from the method's own official framing) designcouncil.org.uk. |
| 6 | Heavy RE standards (IEEE 29148 / BABOK / Volere) | Med | Wins decisively on thoroughness/traceability by design (29148's 9+5 quality-characteristic grid, Volere's 27-section coverage + snow-card fit-criteria) — Tier C/B grounding via standard summaries, not the paywalled primary text — but is the slowest candidate on the list and carries no built-in user-validation loop; justified mainly for complex/regulated builds, which conflicts with "move very fast" for most product work RUP 2026 commentary. |
| 7 | RUP/UP | Low-Med | Same thoroughness strength as RE standards (shared artifact lineage) but adds full phase/discipline process overhead on top — modern consensus (2026) treats RUP-as-process as "largely retired," valuable only as a vocabulary/artifact donor, not as the backbone methodology (Tier C consensus, Wikipedia + practitioner commentary) Rational Unified Process — Wikipedia. |
Reading View 1: no single methodology wins outright — the top 3 win because they are narrow-scope tools that plug into specific stages rather than end-to-end philosophies, which is itself evidence for a hybrid (confirmed rigorously in View 2 below, not asserted here).
3. View 2 — Per-stage fitness ranking#
Each stage ranks all 8 candidates High / Med / Low / N-A (not applicable — method has no defined artifact at this stage per the landscape survey). "N-A" is itself informative: it shows no single methodology spans the whole pipeline, which is the rigorous basis for a hybrid rather than an assumption.
| Stage | RUP/UP | Heavy RE standards | Agile stories/BDD | Lean/CustDev | Continuous discovery/JTBD | Design thinking/Double Diamond | Lightweight tech-spec (C4/ADR/OpenAPI) |
|---|---|---|---|---|---|---|---|
| Discovery / problem-framing | Med (Vision doc input-driven, Tier C) | Low (BABOK has a planning KA but is elicitation-heavy, slow to start, Tier C) | Low (agile has no discovery artifact of its own — borrows from others) | High (Lean Canvas/Customer Development built exactly for this; falsifiable fast, Tier B/C) | High (Opportunity-Solution Tree root outcome + weekly interviews is purpose-built, Tier C) | Med (Empathize/Discover strong on qualitative insight but no completeness gate, Tier A critique on evidence base applies) | N-A |
| Requirements | Med (Use-case model/Supplementary spec, Tier C) | High (29148's measurable quality-characteristic grid, Volere snow cards with fit criteria are the most rigorous, traceable artifacts surveyed — Tier C/B via standard summaries) | Med-High (User story + INVEST + Acceptance Criteria is fast and AI-generatable, but intentionally shallow by design, Tier C) | Low (CustDev produces validated hypotheses, not structured requirements) | Low (JTBD outcome statements feed requirements but aren't requirements themselves) | N-A | N-A |
| UX | N-A | N-A | N-A | Low | Med (Opportunity-Solution Tree extends into UX opportunity-mapping, Tier C) | High (Double Diamond Discover/Define + journey maps/empathy maps/user flows is the only surveyed lineage purpose-built for this stage, Tier A primary source for the phase model, Tier C/🤔 for per-artifact completeness criteria) | N-A |
| UI | N-A | N-A (Supplementary spec Usability category only) | N-A | N-A | N-A | High (Ideate/Prototype/Test ↔ Develop/Deliver, wireframe/prototype fidelity practice, Tier A phase model + Tier B NN/g-adjacent fidelity evidence — noting the headline "5 users find 85%" figure itself has a Tier A caveat: Faulkner 2003 found individual 5-user samples ranged 55-99%) | Low |
| Frontend spec | N-A | N-A | N-A | N-A | N-A | N-A | High (Atomic Design + Storybook + Design Tokens is the only surveyed artifact chain at this stage, Tier C community convention but near-universal tooling adoption) |
| Backend spec | Med (Supplementary spec NFR categories carry over, Tier C) | Med (29148's SyRS/SRS structuring applies but is heavier than needed once use cases exist, Tier C) | Low | N-A | N-A | N-A | High (C4 + OpenAPI + ERD + ADR is purpose-built, machine-checkable, AI-generatable — Tier C practitioner standard but with primary-source grounding for C4, ADR, OpenAPI) |
What View 2 shows rigorously: the winner changes at every stage. Discovery and UX/UI have a clear non-RUP, non-RE-standard winner (lean/continuous-discovery for discovery, design-process lineage for UX/UI). Requirements is the one stage where heavy RE standards genuinely win on the owner's own "maximum thoroughness" criterion — nothing else surveyed has a comparably rigorous completeness test. Frontend/backend spec stages have no competition at all from the discovery/requirements/UX families — only the lightweight tech-spec stack has artifacts there. This is what makes a hybrid derivable, not arbitrary: each stage's winner is determined by "which methodology has ANY defined artifact here, and of those, which has the strongest evidence," not by picking favorites.
4. The speed-vs-thoroughness tension (dedicated, not averaged away)#
The owner's two top-line goals are in real, surfaced tension — not a rhetorical one:
-
At the discovery and UX stages, speed and faithfulness-to-need pull in the SAME direction: Lean Canvas, Customer Development, and continuous discovery are simultaneously the fastest AND the methods explicitly built to prevent building the wrong thing (Ries: work that doesn't produce "validated learning" is waste). Thoroughness in the RE-standards sense (29148, Volere, BABOK) is actively counterproductive here — Tier C: heavy upfront elicitation before any real user contact risks fully specifying the wrong product, precisely the failure mode Customer Development was created to prevent Wikipedia — Customer development.
-
At the requirements and backend-spec stages, the direction flips: thoroughness wins outright and speed is the one taking the hit. Heavy RE standards (IEEE 29148, Volere) and the RUP-descended Supplementary Spec are the only surveyed artifacts with a measurable completeness test (29148's 9+5 quality characteristics; Volere's per-requirement fit criterion). The classic justification for paying this cost is the cost-of-change curve — Boehm's widely-cited claim that a requirements defect caught late costs up to 100x more than one caught early, with IBM Systems Sciences Institute data cited for a "$1 / $10 / $100" design-dev-production multiplier (Tier B, frequently repeated but methodologically thin — original data from 1970s TRW/IBM waterfall projects). This claim itself is now contested: Boehm and Basili's own 2001 revision for small/agile/CI-enabled projects reported a meaningfully flatter curve, and modern commentary argues CI/CD and modular architecture have further flattened it (Tier B/C) Mountain Goat Software — The Cost of Change Curve Is Outdated, reworkcost.com — Boehm Cost of Change Curve: What the 1981 Data Actually Shows. Implication for this tool: the thoroughness payoff at the requirements/backend stage is real but smaller than the classic 100x folklore suggests for an AI-assisted, fast-iteration pipeline — it still justifies rigor at these two stages specifically, but does not justify extending RE-standard-level rigor backward into discovery/UX, where it would slow exactly the stages that most benefit from speed.
-
Net resolution (not an average, a per-stage split): the tension is not resolved by picking one methodology-family for the whole pipeline — it is resolved by letting the winning methodology change per stage (View 2), with thoroughness concentrated specifically at requirements and backend-spec (where it is cheap to be thorough and expensive not to be) and speed/faithfulness concentrated at discovery/UX (where thoroughness without user contact risks building the wrong thing well). This mirrors the "two configurable depths per stage" accommodation already identified in
stage-chain-matrix.md's Ordering Disagreements section — this ranking confirms, with evidence tiers attached, that the accommodation is justified rather than a convenient compromise.
5. Evidence-tier honesty check (escalation rule)#
Tally of what this ranking actually rests on:
- Tier A (peer-reviewed / large-N): Faulkner (2003) on usability-sample variance; Hobday/Boddington/Grantham critique of design thinking's evidence base; Serrador & Pinto-style agile/waterfall comparative study (via secondary citation — primary not independently refetched, flagged); the Double Diamond / d.school phase models as officially-sourced primary frameworks (not effectiveness studies, but primary methodology sources).
- Tier B (industry reports, flagged for known criticism): Standish CHAOS-derived agile-vs-waterfall success-rate figures (criticized for self-selected samples and a success definition that changed in 2015 from strict scope/cost/time to a looser "satisfactory result" test — reported here with that caveat attached, not averaged away); Ulwick/Strategyn's 86%-vs-17% ODI claim (single-vendor, not independently replicated); Boehm/IBM cost-of-change figures (1970s waterfall-era data, now contested by a 2001 revision and modern CI/CD commentary).
- Tier C (practitioner consensus, no controlled study): RUP's 2026 retirement-as-process consensus; Storybook/Atomic Design completeness convention; C4/ADR/OpenAPI practitioner-standard status; Volere/29148/BABOK completeness-grid descriptions (grounded via secondary summaries of standards whose primary text is paywalled).
Rough mix: roughly 25% Tier A, 40% Tier B, 35% Tier C across all scores in this document — this is NOT over the 50%-Tier-C escalation threshold, so the overall ranking (View 1) is retained alongside the per-stage view, not dropped. The escalation rule is noted as a close call, not triggered: the two scores leaning hardest on Tier C alone (RUP's "largely retired" status, and Frontend-spec's "Storybook convention" win) are both low-controversy claims where practitioner consensus is unusually strong and uncontested, which is why they are kept at Tier C rather than treated as weak points requiring a hedge.
Sources (new to this document, beyond methodology-landscape.md)#
- Assessing IT Project Success: Perception vs. Reality — ACM Queue
- Standish CHAOS Report — definitional change 1994→2015, budgetoverrun.com
- Waterfall and Agile information system project success rates — South African perspective, scielo.org.za
- Project Success in Agile Development Projects — arXiv
- A Systematic Mapping Study on Software Engineering Practices to Develop MVPs — arXiv
- Jobs-to-be-Done / ODI 86% success claim — Strategyn
- The craze for design thinking: a critique — ResearchGate (Hobday, Boddington, Grantham)
- The "magic number 5": is it enough for web testing? — ResearchGate (Faulkner 2003 discussion)
- The Cost of Change Curve Is Outdated — Mountain Goat Software
- Boehm Cost of Change Curve: What the 1981 Data Actually Shows — reworkcost.com
- Rational Unified Process — Wikipedia
- UK Design Council — The Double Diamond
Grounding caveats specific to this ranking#
- Serrador & Pinto's exact "16% higher success rate" figure and the South African study's 88.2%/47% figures were relayed via aggregated search-result summaries, not independently re-fetched from the primary paper PDFs — treated as Tier A-sourced-via-secondary-aggregation, one level removed from the primary text, consistent with how methodology-landscape.md already treats similarly-sourced claims (e.g. its Scrum Guide DoD caveat).
- The CHAOS-derived "agile more successful than waterfall" headline is reported together with its own criticism (self-selected sample, no disclosed methodology, shifting success definition) per this task's explicit instruction — never presented as a clean Tier A fact.
- Ulwick's 86%-vs-17% ODI success claim comes from Strategyn, the commercial originator of the ODI method — a single, non-independent source; treated as Tier B with a conflict-of-interest flag, not elevated.
- No controlled, independent effectiveness study for Teresa Torres's Continuous Discovery Habits was located in this pass (only book summaries and the author's own site); the "High" score for continuous discovery in View 1 rests on structural fit-to-goal reasoning (weekly real-evidence cadence directly answers "exactly what the user wants") rather than a proven-effective claim — flagged as the weakest-evidenced "High" score in this document.