What would actually falsify this
An earlier version of this site offered an internal test: model an organisation in which adding an approval layer improves structural health and the framework is wrong. That test cannot come out false, because Decitect's own scoring function decides whether the layer improved anything. It is a consistency check wearing a falsifiability costume; it has been retired.
The real test is external and blind:
- Assemble a set of real organisations whose structural outcomes are independently documented: sustained delivery at scale on one side, delivery collapse or a structural post-mortem on the other. The outcome record is fixed before anything is scored.
- A modeller who does not know the outcomes builds each organisation's Decitect model from structural facts alone: the org chart, the dependency map and the authority placement as they stood at the time.
- Score each model once. No re-modelling after seeing a result.
- Test whether the score ranking predicts the outcome ranking better than chance, against acceptance thresholds fixed before any score is computed.
Fixed where matters. A threshold stated only on this page is timestamped by the same person who computes the scores, which a reviewer can discount for free. The full protocol, thresholds included, therefore lives in the repository as PREREGISTRATION.md, written for deposit with an independent registry (OSF) whose timestamp is not the author's to edit. Until that deposit exists, the honest reading is that the bar is public but self-timestamped; discount accordingly.
This protocol has not been run. The obstacle is data honesty rather than design: structural facts about failed organisations are confidential and post-hoc accounts of them are unreliable, so assembling a clean blind set is slow. Until it runs, the correct status of the model is: a published, inspectable expert prior whose external validation is specified but outstanding. If you hold structural records that could join such a set, the repository is the place to say so.
Sensitivity: do the conclusions depend on the exact numbers?
If the published results only hold at precisely these coefficients, the model is a curve-fit to its own examples. So the repository carries a deterministic sweep, sensitivity.py, which scales every tunable coefficient by 0.8 and by 1.2 (the three composite penalty shares are renormalised to keep their enforced sum of 1.0; the prince band's float coefficients are swept while its headcount edges, being structural counts, are not), re-scores the ten shipped archetypes under each of the 36 perturbed configurations and checks five qualitative conclusions: the typical archetypes still collapse in order with scale, the well-designed archetypes keep their order, every well-designed archetype still outscores its typical counterpart and both canonical blunders (the approval layer and the matrix overlay) still lower the score on every typical archetype.
All five conclusions survive every perturbation. The baseline scores below reproduce the published archetype numbers exactly; under all 36 perturbed configurations the orderings, the typical-versus-well-designed gap and the signs of both blunders are unchanged. The qualitative results come from the structure of the model, not from the particular values of its coefficients.
An axis sweep only explores the edges of the box; models typically break in the interior, where several coefficients drift at once. So the same script runs a second, joint sweep: a Latin hypercube of 1,000 draws inside the same ±20% band with every coefficient moving simultaneously, each draw renormalised and capped so it respects the model's own validation, reporting the fraction of draws under which each conclusion holds. All five conclusions hold in 100% of the 1,000 joint draws (seed 20260711, the same published seed the books use).
The sweep has already earned its keep once: the first cut of contested ownership counted a team's own authority as its only structural claimant; under that rule the matrix overlay read positive on the escalation-heavy archetypes (it barely contested anyone while diluting every penalty share). The sweep failed, the semantics were corrected (the structural owner is always claimant one, so any standing claim is contest) and the failure is recorded here rather than erased.
| Archetype | Typical | Well-designed |
|---|---|---|
| Startup | 83.4 | 98.7 |
| Scale-up | 44.9 | 88.9 |
| Enterprise | 18.8 | 75.9 |
| Very large | 15.1 | 66.8 |
| Conglomerate | 12.0 | 64.9 |
The sweep's own limits, stated plainly: it stays within ±20%, against the ten
shipped archetypes. It does not test larger swings or every published claim (the
1,732-team move-locality result is not re-run under perturbation). And the archetypes
themselves are authored by the same hand as the coefficients: both sweeps test
robustness to the numbers while holding fixed test cases that were built to exhibit the
very properties the model penalises. It is a robustness check on the published
conclusions, not a proof about every organisation. Rerun it from the repository root
with python sensitivity.py; the axis sweep has no randomness and the joint
sweep draws from the published seed, so you will get these numbers exactly.
Decitect