Open the methodology changelog
# Clavis Rating Methodology Changelog
All notable changes to the Clavis Rating methodology are documented here. Versions follow semantic-style conventions:
- **Major** (vN.0.0): pillar added or removed, composite formula reshaped
- **Minor** (v0.N.0): pillar calculation changes meaningfully
- **Patch** (v0.0.N): parameter tuning within an existing calculation
Every change references the spec doc (`vN.N.N.md`) and YAML (`vN.N.N.yaml`) that defines it.
---
## Public rating gate: no letter where the state publishes no severity (coord implementation, 2026-09-29)
Not a methodology version. No pillar, weight or formula changes, and no
published California number changes. This is a coord implementation of
the fail-closed principle already ratified with the Thread 127
confidence gate (a letter is published only where the data can defend
it), pending Scott's separate decision on how Texas is ultimately
presented.
The gap: Thread 129 withholds the letter for a facility whose citation
rows carry an ingest's "severity not published by the state" marker.
A marker lives on a citation, so a Texas facility whose inspection rows
landed with no citations attached carried none and would have published
a letter on the California scale. On the 2026-09-07 roster that was
1,908 of 2,103 Texas facilities (`ops/COORD_HANDOFF.md`, "Live trap";
`ops/STATE_DATA_GAPS.md` TX-1 to TX-3).
The rule: `clearview/severity_confidence.py` gains
`SEVERITY_PUBLISHING_STATES`, an allowlist holding California only, and
`clearview/rating_confidence.py::rating_status` takes a required
`state_publishes_severity` argument. A facility with an inspection record
in a state not on the list returns `withheld_unpublished_severity` (the
outcome Thread 129 introduced, reused because the fact is the same) and
the API sends no score. Skilled nursing facilities are graded on the CMS
A-to-L letter and are not affected. A facility with no inspection record
still reads `pending_inspection_data`. The list only withholds; it is
checked in addition to the marker, never instead of it.
Consequence for Thread 129's design: a state no longer clears the gate by
its re-ingest dropping the marker. Releasing a state is a one-line change
to the allowlist with a test beside it. Florida is left off because
`ingest_fl.py` stores STANDARD without a marker when it finds no class.
Evidence: `backend/tests/test_no_severity_state_gate.py`, including a
comparison of every California input against the predicate as it stood
before. The community page carries copy for the outcome
(`frontend/src/lib/rating.ts`, `ratingNoSeverityLines`).
---
## Risk Index v0.2.1, Section A CMS window — ratified 2026-09-07
The displayed Section A citation metrics (`sev_rate_100beds`, its national
and state percentiles, `pts_cycle1..3`, `trend`, and the `risk_index_draft`
blend that reads them) now count, for certified nursing facilities as for
assisted living, only citations dated in the three years to the as-of date,
cut into three twelve-month periods counted back from it. Previously the CMS
path counted every citation CMS stamped with a survey cycle, with no date
test, and divided by three on the assumption that one cycle approximates one
year; on the 2026-08-26 release the median facility's three cycles covered
4.16 years. The divisor of three is now true by construction, and the
displayed rate equals the calibrated model's `sev_rate_100beds_3yr`.
Measurement and options: `research/calibration/cms_cycle_divisor_2026-09.md`
(Option 3, §7).
Approved by the founder 2026-09-07 (`ops/COORD_HANDOFF.md`, "Decisions Scott
made on 2026-09-07", item 1). This is an editorial choice and the method note
says so: a citation older than three years drops out of the displayed rate
even when CMS still publishes it (84,089 citations, 20.1% of the counted set,
at as-of 2026-06-30), and a facility with no citation inside the window shows
a rate of zero beside the stale-survey flag. The memo's measured effect: the
national median severity rate moves from 18.33 to 14.05 and the Ensign median
from 19.285 to 15.74. Committed demo artifacts under `research/demos/` were
computed on the old definition until regenerated. Section B features are
unchanged. Spec doc amended in place (`v0.2.0.md` §A rows for the rate and
trend).
---
## Risk Index v0.2.1 — ratified 2026-09-03
Amendment to the professional Risk Index v0.2.0 (parallel to the public
rating; does not touch v0.1.0). State residual becomes a reported
diagnostic rather than a gate; Section B excludes state intercepts and
state outcome-rate terms from the facility score; a provisional
calibration tier is added below the 0.70 production bar with mandatory
disclosures. Proposal and evidence: `research/methodology/v0.2.1_proposal.md`.
Ratified 2026-09-03 on coord's recommendation under AGENTS.md §9;
revertible by the founder. Option (b) with a soft (c) threshold and the
"calibrated, provisional" tier, explicitly not (d) or (e). Spec doc and
YAML amended in place (`v0.2.0.md`, `v0.2.0.yaml`; the file names stay).
The amendment stages a number, it does not render one. `v0.2.0.yaml`
gains `active_index`, which reads `draft` and keeps the Section A
percentile blend as what deliverables show. Flipping it to `provisional`
is the founder's decision and is the only change that alters a rendered
number.
---
## v0.1.0 — 2026-06-01 (ready for shadow-score evaluation)
**Reframe from deduction model to three-pillar composite. Implemented end-to-end.**
### Implementation
- `clearview/scoring/v01.py` — three-pillar Safety/Trajectory/Sentiment
computation with acuity, density, and vintage corrections.
- `clearview/scoring/sentiment.py` — Sentiment LLM extraction interface
+ recency-weighted aggregation + DB persistence.
- `clearview/scoring/shadow.py` — shadow-score diff tool. Compares any
two methodology versions across the full facility corpus and emits
a markdown distribution/migration report.
- `clearview/scoring/tinker.py` — CLI for interactive parameter
experimentation against real or synthetic facilities.
- `backend/routes/methodology.py` — `/api/methodology` endpoints
exposing the active methodology, spec doc, and changelog.
- `frontend/src/app/methodology/page.tsx` — public methodology page.
- `clearview/scripts/calibrate_acuity_benchmarks.py` — calibration
tool for the acuity benchmarks; placeholders shipped pending
real-corpus calibration.
### Schema (migration d9e0f1a2b3c4)
- `facilities.year_built` (Integer, nullable) — vintage-control input.
- `facility_scores` extended with per-pillar columns plus adjustment
flags (`safety_acuity_adjusted`, `safety_density_normalized`,
`safety_vintage_adjusted`), `insufficient_data` gate flag,
`care_type_used`, and Trajectory/Sentiment sub-component scores.
- New `community_sentiment_scores` table — per-dimension Sentiment
scores keyed on facility + dimension + methodology version.
### Launch-posture decisions documented in spec
- **Publication gate relaxed:** `require_at_least_one_non_safety_pillar`
set to `false` for v0.1.0 launch because Sentiment data does not yet
exist for the corpus. Safety alone is publishable. Flip back to
`true` in v0.1.1 once Sentiment is populated for the majority of
facilities.
- **Acuity benchmarks shipped as placeholders** (IL 0.04, AL 0.10,
MC 0.18, SNF 0.31) pending real-corpus calibration via the
`calibrate_acuity_benchmarks` script. Recalibration will be a
v0.1.1 patch.
### Added
- Three-pillar composite: Safety, Trajectory, Sentiment.
- Care-type-specific composite weights (IL, AL, MC, SNF).
- Care-type-specific deficiency benchmarks calibrated from CA 2023-2025 corpus.
- Inspection-density correction to normalize for state regulator inspection cadence differences.
- Vintage control on physical-plant deficiencies for buildings older than 10 years.
- Time-to-correct sub-component (median days from `date_found` to `date_corrected`).
- Repeat-violation-rate sub-component (same regulation cited within 18mo prior).
- Volume-gated dimensional Sentiment scoring (LLM extraction from Google reviews on 8 dimensions).
- Care-type-specific dimension weights inside Sentiment.
- Confidence bands and publication rules for sparse-data communities.
- Weight redistribution rules when Trajectory or Sentiment data is insufficient.
- `methodology_version` stamp on every score row.
### Changed
- Severity, Frequency, Recency, Complaints, Inspections (the v0.0.0 components) are now sub-components inside the Safety pillar rather than direct contributors to the overall score.
- Frequency benchmark is now care-type-specific instead of universal 0.10/bed.
- Trajectory is no longer change-in-severity (which double-counted Safety). It is now time-to-correct + repeat-rate, genuinely independent of absolute deficiency count.
### Removed
- Universal deficiency-rate benchmark (replaced by care-type benchmarks).
- Implicit assumption that all deficiencies are equally weighted regardless of building age (vintage control now applied to physical-plant subset).
### Migration notes
- All existing facility_scores rows stay valid with `methodology_version = "v0.0.0"` backfilled.
- v0.0.0 scores remain queryable for historical comparison.
- Shadow-score evaluation window: 14 days of dual-computation before v0.1.0 becomes the published source of truth.
### Open questions deferred to v0.2.0+
- Stratified Safety scoring (per care type within mixed buildings).
- Renovation events in vintage control.
- Operator-level Clavis Rating as a parallel product.
- CMS Care Compare integration → SNF Outcomes pillar.
- Sentiment confidence intervals on volume.
- EHR-driven Outcomes pillar (long horizon).
- Pricing transparency / Access pillar (when rate-card coverage broadens).
---
## v0.0.0 — production through 2026-05-25
**The deduction model. Current production.**
Five components summed with fixed weights:
```
Overall = 0.30 × Severity
+ 0.20 × Frequency
+ 0.20 × Recency
+ 0.15 × Complaints
+ 0.15 × Inspections
```
Universal deficiency-rate benchmark (0.10/bed). No acuity adjustment. No inspection-density correction. No vintage control. No trajectory signal. No consumer sentiment input. Same calculation across IL/AL/MC/SNF.
Implementation lives in `clearview/scoring/engine.py` and `clearview/scoring/config.py`. Preserved unchanged for historical comparison. New score rows after v0.1.0 launch use the new methodology; old rows retain the v0.0.0 stamp.