Version: April 2026
Berkeley faces real governing problems: fiscal strain, infrastructure backlog, pension costs, and public safety pressure. City Council has limited time and resources. These scorecards evaluate whether members use them to deliver real results for Berkeley residents.
Berkeley faces real governing problems. The City Council oversees a mid-sized city with a $630M budget, a structural fiscal deficit city staff has described as unsustainable, $2.1B in unfunded capital obligations and deferred maintenance, $695M in net pension liability, persistent pressure on public safety, and labor agreements that constrain what any of those problems can look like as policy. Managing those realities requires discipline, prioritization, and sustained attention. A Strategy for Strategy, the budget framework published alongside these scorecards, is a proposal for supplying exactly that — a repeatable test applied before the politics of any individual proposal begin, sorting what the city must ensure, what it should undertake only with committed partners, and what it should not own at all.
But the incentives of local politics do not always reward those qualities, and several of the recurring decision problems documented on the Berkeley Decision Project are what that looks like in practice. A bond measure gives a member something to campaign on; holding a recurring budget line steady does not — Berkeley ran five street measures while the recurring allocation sat unchanged since 2014 (Bonds Win Elections). Buying a building is a ribbon-cutting; the thirty-year operating bill arrives long after (Buy Once, Maintain Forever). Deliberation fires on contention rather than magnitude, so commitments the city carries for decades pass unremarked while a symbolic resolution consumes an evening (Two Calendars). And a consultation narrowed before it reaches residents returns support rather than an inconvenient answer (Which Version, Never Whether).
The same asymmetry explains the option set itself. Raising revenue distributes a cost thinly across taxpayers who are not in the room; reprioritizing, finding efficiencies, or asking whether someone else could deliver a service better all impose a concentrated cost on an identifiable group that is — staff, a department, a constituency, an organized bargaining unit. The first requires no one to be told no. The others require naming what stops, and naming it to people who will be at the next meeting. That is a structural reason the four money-raising options stay in use and the three that would change what the city spends do not (The Crayon Box), and it holds without attributing any motive to any member.
A June 2026 item shows the shape of it. Berkeley's environmental health division does not cover its costs; even after a fee increase, a councilmember put the remaining General Fund subsidy at roughly $750,000 a year on the record, uncontested. Alameda County performs the same inspections for other cities — Berkeley, staff told Council, is "one of the few cities in the state that has its own environmental health division." Two councilmembers asked what transferring the function would look like. Staff answered that they could analyze the costs and benefits, including "the benefits of having staff who can do this work," and the Council adopted the fee increase unanimously. No referral was made and no analysis was directed. The benefit of keeping the work in-house was asserted; the cost of doing so was already quantified.
The counterfactual is worth stating plainly, because it is the part of this that never gets said out loud: a member who moved to consolidate a function, contract a service out, or eliminate a class of position would be imposing a concentrated, visible loss on an organized constituency that turns out, endorses, and will be at the next meeting. Would they expect that constituency's support at the next election? The question answers itself, and it explains why an option can remain formally available and practically unused. Nothing here suggests any member weighed it that way. The point is that no one has to.
The constraint is visible in how efficiency gets proposed when it is proposed. In July 2026 a member raised piloting automation on the City's after-hours answering service — and noted, before making the case, that the service "is an outsourced service. It does not affect any city staff." The reassurance was offered unprompted, about a contract already held outside the workforce. That is what an unspoken boundary looks like from the inside: not a rule anyone states, but a line proposals are shaped to avoid crossing.
Not every habit has a reward behind it. Some are failures of capability rather than incentive — a city that cannot measure whether its programs work has no basis on which to rank them, and nobody gains politically from that. The two compound each other: without outcome data, a member who wants to stop something cannot show it is failing, so the concentrated cost of proposing it lands with nothing to justify it. The distinction still matters for scoring, though. These scorecards assess conduct against a documented standard, not motive — a member is credited for asking the question and putting an alternative on the table, whatever the reason it was hard to do.
Berkeley residents also live under real financial constraints of their own. Housing costs are high. Everyday expenses are high. Many households have limited tolerance for continual tax increases, new parcel taxes, higher fees, or additional debt layered onto an already expensive city. Responsible governance must account for the fiscal capacity of residents as well as the fiscal needs of government.
Council members are often rewarded for highly visible constituent casework, ideological signaling, symbolic resolutions, and commentary on issues beyond the council's practical authority. Those activities may generate attention or political goodwill, but they do not necessarily address Berkeley's core fiscal, operational, and governance challenges.
The record bears this out in the plainest possible terms. A resolution on a foreign ceasefire drew 197 speakers and a three-hour public comment period extended twice; a single street's paving alignment drew 160. In the same record, the item establishing whether the city can measure if any of its programs work passed on the consent calendar with no debate, no speakers and no letters — see Measuring Nothing. The council is entirely capable of sustained deliberation. It allocates that capacity to contention rather than to consequence.
That tension is the reason these scorecards exist.
Berkeley residents are asked to evaluate elected officials in an environment that is difficult for any ordinary voter to track. Council meetings often run four to five hours. Agendas can span hundreds of pages. Staff reports, audits, budget documents, commission materials, newsletters, and public statements accumulate continuously.
No resident with a normal life can reasonably monitor all of it.
The common substitutes are also limited:
What is often missing is a sustained evaluation of how officials actually use the office over time — whether their words and deeds part company, how much time goes to matters that don't address Berkeley's documented problems, and whether they take ownership of hard issues or deflect.
These scorecards are an attempt to provide that evaluation. They examine the public record, including:
The purpose is not to reward charisma, popularity, or ideological alignment. It is to assess governing performance.
These scorecards use an explicit and transparent governing standard: elected officials should devote the council's limited time and the city's finite resources to Berkeley's most pressing documented problems.
That includes matters such as: fiscal sustainability, infrastructure and maintenance backlogs, public safety performance, housing and land-use execution within city authority, service delivery and administrative competence, and long-term stewardship.
Officials who focus on those responsibilities, make tradeoffs honestly, and help move solutions forward receive credit. The scoring cuts both ways: time and votes directed at matters outside Berkeley's jurisdiction count against the score, and documented local crises that go unaddressed count against it too. Being excellent at things a city council cannot change does not offset failing at things it can.
These scorecards do not claim to be value-free. Every political evaluation rests on assumptions, whether disclosed or hidden.
The assumptions here are stated openly: municipal office should primarily be used to govern the municipality well.
Reasonable people may disagree with that framework, the weighting of criteria, or specific judgments. That is legitimate. The methodology is published so those disagreements can be concrete and substantive rather than implicit.
Healthy democratic accountability requires more than elections every few years. It requires understandable records, coherent standards, and public argument about performance.
These scorecards are offered in that spirit — not as the final word, but as a serious attempt to ask and answer an important question: how well is Berkeley governed, and who is helping govern it well?
This scorecard is explicitly voter-aligned, not neutral. The evaluative framework assumes the voter cares about:
Taxpayer alignment — Does the member champion taxpayer interests, or treat property owners and residents as a funding source for their agenda? Do they demand alternatives to taxes and bonds, question efficiency, and push back on the status quo? Or do they reach for new revenue as a first resort?
Focus — Does the member spend the council's time on core city services (public safety, infrastructure, basic city operations), or on performative, ideological, and non-core items that consume staff bandwidth and budget without delivering core value?
Attendance & Vote Presence — Did the member actually do the job? Attendance at meetings — especially for binding fiscal votes — is the minimum bar.
This is not a promise-keeping scorecard. A member who campaigned on housing affordability and never mentioned the structural deficit is not evaluated on housing affordability. They are evaluated on whether they engage with Berkeley's documented Priority 1 (P1) fiscal crises — because those crises exist regardless of what any member promised, and a representative who ignores them is not serving the taxpayer regardless of their campaign platform. A voter who wants to build a scorecard measuring housing production or homeless services expansion can do so; this one measures something different.
The scoring does not treat budget growth, new programs, or bond issuances as neutral acts. A YES vote on a budget adoption is a choice to endorse the status quo and forgo reprioritization. An absence during a major fiscal vote is a failure of the core duty of the office. A referral to study a new tax is the beginning of a political infrastructure campaign, not a neutral process step.
Each member scorecard is divided into two explicit layers:
Layer A — What the Record Shows: Objective measures drawn from official records and attributed transcript speech. No weighting or judgment applied. Includes: attendance, major fiscal vote participation, voting record (NO votes, abstentions, absences), items authored/cosponsored, staff direction volume, core/non-core speech share, fiscal concern mentions, new revenue preference mentions, spending votes. These facts are nearly unimpeachable — critics who dispute the overall grade must engage at the layer B level (philosophy and weights), not at this level.
Layer B — Fiscal Stewardship Assessment: The scorecard's normative framework applied to the layer A facts. Includes: Homeless Services Status-Quo Alignment (HSA) score, Taxpayer Alignment breakdown, composite grades and pillar scores, key findings. This layer explicitly encodes the scorecard's values: taxpayer alignment, focus on core services, attendance accountability. Reasonable people can disagree with these weights. The structure makes that disagreement productive — it is a dispute about the normative layer, not about the facts.
| Source | What it provides |
|---|---|
| Meeting transcripts (51 PDFs, txt) | Attributed speech per member: rhetoric, debate style, fiscal language, special interest alignment signals |
| Annotated agenda PDFs (58 sessions) | Authoritative post-meeting record: exact attendance (roll call time, arrivals), per-item vote breakdown (Ayes/Noes/Abstain/Absent) |
| Agenda JSONs (58 sessions) | Pre-meeting agenda items: title, dollar amount, authors, cosponsors, non-core flag, fiscal flags |
| Staff report PDFs (packet scraper) | Procurement signals: waived competitive bid, backdated contracts, no-alternatives clauses |
| Budgets and CAFRs (Finance Director) | Adopted annual budgets and Comprehensive Annual Financial Reports establish the city's own financial self-description: fund balances, appropriations, staffing levels, debt obligations, and the City Manager's budget message language. Distinct from City Auditor reports in origin and purpose: CAFRs are management-prepared financial statements with an external audit opinion on fair presentation; City Auditor reports are independent performance and operational audits initiated by the auditor, not management. Used to establish factual baseline — revenue trends, expenditure growth, reserve levels — against which council behavior is evaluated. |
City Auditor reports (audit_findings.json) |
Independent documented findings that establish ground truth: what the city's financial condition actually is, what structural problems exist, and what warranted action looks like. Audits are tracked in a separate registry; the council's response — or non-response — is the scored event. Three reports currently tracked: Rocky Road streets audit (Oct 2025), Homeless Response Team audit (Jul 2025), Financial Condition audit (Apr 2026). |
Incidents (incidents.json) |
Out-of-meeting behaviors: constituent interactions, public statements, newsletters, and patterns not captured in formal proceedings. Each incident is assigned one of three evidence tiers — A (primary public record), B (reputable reporting or member communications), C (direct observation) — which determine its effective scoring weight. Incidents with an audit_ref field are grounded in a specific City Auditor finding. |
Annotated agendas are the authoritative source for what actually happened (outcomes, votes, attendance). Transcripts fill the gap between plan and outcome — capturing the deliberation, rhetoric, and interpersonal style that votes alone don't reveal. City Auditor reports provide the independent factual baseline against which council action is evaluated.
Why this reference standard matters: The Priority 1 (P1) problems scored here are grounded in the city's own professional documents — consecutive City Manager budget messages, City Auditor findings, infrastructure condition reports — not in the scorecard author's political preferences. A member who disagrees with the evaluation cannot simply claim it is ideological; they must dispute the City Manager's finding that the structural deficit is "not sustainable," or the Auditor's finding of a 66% pension funded ratio, or the MTC's finding of PCI 57. The city's own organization has documented what is broken. The scorecard asks whether elected officials are engaging with those findings.
A single A–F composite representing overall alignment with taxpayer interests. Computed from weighted sub-scores (see below). Designed to be the one number a busy voter can reference.
Composite formula:
composite = max(0.0,
taxpayer_alignment × 0.55 + focus × 0.25
+ lsi × 0.10 + character × 0.10
− attendance_deduction
− low_engagement_adj
+ audit_penalties
)
Where taxpayer_alignment = taxpayer_base × 0.60 + audit_composite × 0.40.
Component weights and deductions:
| Component | Role | Notes |
|---|---|---|
| Taxpayer alignment | 55% weight | The core question: whose interests does this member champion? (blended with audit alignment at 60/40) |
| Focus | 25% weight | What did they spend the council's time on? |
| LSI | 10% weight | Legislative sophistication (domain knowledge, fiscal literacy, inquiry rate, decisiveness, process) |
| Character | 10% weight | Collegiality, humility, warmth, low self-referential appeals |
| Attendance deduction | Up to −0.30 | Convex curve: lenient for 1–2 absences, severe for 4–5 (see Fiscal Vote Record) |
| Low engagement adjustment | Up to −0.10 | Triggered when member authors no Priority 1 (P1) referrals AND shows low fiscal engagement (see P1/P2/P3 Framework) |
| Audit-grounded penalties | −0.05 to −0.20 | Applied after composite blend: structural silence, one-time masking, cross-fund transfers, Section 115 depletion |
Grade thresholds: A+ ≥ 90 · A ≥ 83 · A− ≥ 77 · B+ ≥ 70 · B ≥ 63 · B− ≥ 57 · C+ ≥ 50 · C ≥ 43 · C− ≥ 37 · D+ ≥ 30 · D ≥ 23 · D− ≥ 17 · F < 17
Rhetoric and behavior are time-sensitive. A member who said concerning things two years ago but has since moderated should not carry that signal at full weight indefinitely. Conversely, a member who has been consistent over the entire period should score the same as one who is only recently reform-oriented.
Formula: Each time-sensitive signal is weighted by:
effective_weight = e^(−λ × age_in_years) λ = 0.7
| Age | Weight |
|---|---|
| Current | 1.00 |
| 1 year | ~0.50 |
| 2 years | ~0.25 |
| 3 years | ~0.12 |
Applies to:
- Incident scoring (each incident's scoring_impact × tier_weight is multiplied by the decay factor before summing)
- Rhetoric signals: fiscal concern rate, new revenue preference rate, self-referential appeals, collegiality, humility, warmth (via per-meeting decay-weighted aggregation)
- Newsletter Priority 1 (P1) silence events
Does NOT apply to:
- Bond/tax referral votes (durable official acts — authoring a ballot measure is a permanent record)
- Annotated agenda votes (YES/NO/absent on spending items)
- Major fiscal votes
- Homeless Services Status-Quo Alignment (HSA) score (full-text aggregation without per-meeting dates; planned for future update)
Opt-out: Individual incidents can carry "no_decay": true to exempt them from decay (e.g., a member who authored a bond measure that is still active).
Each member scorecard and the summary document include a ranked list of the highest-leverage actions a member could take to improve their composite score. This section explains the logic behind those recommendations.
Five opportunity types are evaluated, ranked by estimated composite impact, and the top five are shown:
| Opportunity | Trigger condition | How it helps |
|---|---|---|
| Demand HSA outcome metrics | Homeless Services Status-Quo Alignment (HSA) score ≥ 45 | Shifts HSA rhetoric toward accountability-oriented language |
| Back fiscal concerns with votes | High concern rate, zero NO votes, and HSA ≥ 45 or ≥ 3 fiscal vote absences | Closes rhetoric–action gap; strengthens composite_fiscal_ref_penalty |
| Attend major fiscal votes | fiscal_vote_absent > 0 | Eliminates attendance deduction |
| Redirect non-core authorship | composite_off_penalty > 0.03 | Reduces scope penalty in Fiscal Stewardship Alignment |
| Pair revenue advocacy with reprioritization analysis | new_revenue_preference_rate ≥ 0.3 | Reduces new revenue preference penalty |
| Engage with audit findings | audit_alignment_composite < 0.35 | Improves audit alignment sub-score |
| Increase fiscal engagement | fiscal_raw < 0.15 (fallback) | Improves fiscal discipline component |
Estimated impact is expressed as composite grade points (0.0–1.0 scale). Items are sorted by descending estimated impact; ties are broken by definition order.
These are directional recommendations — they identify where the data shows the largest gaps. They do not guarantee a specific grade improvement because other signals could change in the same direction.
The summary PDF includes a separate page of systemic findings that no single member can fix alone. These reflect council-wide norms, agenda management practices, and collective voting behavior:
| Finding | Trigger |
|---|---|
| Reduce bloc voting | Block vote rate ≥ 60% |
| Fiscal pre-vote questions norm | ≥ half of members with fiscal_raw < 0.30 |
| Triage staff referrals | Total referrals across all members ≥ 5 |
| Rhetoric–action gap | ≥ 2 members with rhetoric_action_gap_score ≥ 0.15 |
| Refocus non-core authorship | ≥ 3 members with composite_off_penalty > 0.05 |
| HSA accountability metrics | ≥ 55% of members with hsa_score ≥ 55 |
Council-wide findings are not scored — they document structural patterns for voters and elected officials who want to change those patterns.
A+ is not available to a member who is silent or miscalibrated on Berkeley's documented structural crises. The grade requires not just absence of bad behavior but aspirations set at the level the city's own documents say the problems demand. Setting sights too low on a documented crisis accepts a trajectory the city's own adopted reports describe as ruinous.
The core test applies to any domain where significant public money is spent:
A serious elected official demands measurable outcomes for significant public expenditures. This applies equally to street maintenance, homeless services, public health, contracted social services, and staffing. The amount spent and the absence of measurement are what trigger the standard, not the political valence of the program. A member who approves or tolerates large recurring spending in any domain without requiring performance metrics, outcome data, or accountability mechanisms is not doing the job.
Silence is not neutrality. A member who does not demand accountability for a major spending program is implicitly endorsing it.
HSA is one instance of the general principle — the one with the richest transcript signal and the best-documented evidence of unaccountable spending. Berkeley's homeless services apparatus ($21M+/yr, 33+ programs, Housing First mandate, decade of growth with no measurable reduction in visible homelessness) is where the scoring algorithm is best instrumented, not because homelessness is uniquely important but because the evidence record is clearest. A separate signal would be built for any other major program that shares the same profile: large recurring cost, clear measurable objective, council has not demanded outcome data.
HSA measures how invested a member is in the existing apparatus versus demanding accountability, outcome metrics, and reform. Scale: 0 (reform-oriented) → 100 (status-quo aligned). The scoring formula uses a quadratic curve: HSA 50 (neutral/silent) → taxpayer_alignment contribution capped at 0.25 of its maximum. A+ requires HSA well below 50 — not because demanding accountability for homeless services is uniquely virtuous, but because this is the largest unaccountable program in Berkeley's budget and neutrality on it signals a broader tolerance for unaccountable spending.
The following are known gaps where the A+ ceiling should apply but the algorithm does not yet detect it. Each follows the same structure as HSA: significant recurring cost, measurable standard the city's own documents establish, council response that falls short.
Infrastructure outcomes accountability. The city's adopted goal is PCI 70. The Rocky Road audit (Oct 2025) found: current PCI ≈ 57; achieving PCI 70 requires $42M/year; the GF allocation is $2–8M/year; deferred maintenance costs $7 in repairs for every $1 spent early; delays will quadruple costs by 2050. PCI 70 is not an aspirational ceiling — it is the floor below which deterioration accelerates rapidly. A member who accepts PCI 70 as the goal without demanding the $42M/year to get there has accepted a 5× funding shortfall as normal. Same structure as HSA: large program, measurable objective, council has not demanded the outcome. Gap: transcript signal cannot yet distinguish whether a member treats PCI 70 as a ceiling or a floor.
Structural balance policy gap. The City Auditor explicitly recommended the council adopt a GFOA policy requiring assessment of whether recurring revenues match recurring expenditures. No member has moved to adopt it. Three consecutive budget cycles have included verbatim "not sustainable" findings. A council that receives that language repeatedly without adopting a structural balance requirement is not demanding accountability from its own budget process. Gap: motion has not yet been made; absence of motion is the finding.
Reserve policy backslide. The reserve target was lowered from 30% to 20–30% in July 2025, timed to avoid non-compliance rather than achieve the original goal. Combined reserves are ~14.5%, below even the revised floor. An A+ member voted against the revision or demanded a plan to reach the original target. Gap: vote is on record; not yet wired into scoring.
Investment policy non-compliance. City investments have underperformed LAIF for 9 consecutive quarters. The council receives quarterly evidence and has not acted. An A+ member has demanded corrective action on the record. Gap: quarterly reports are on the action calendar; member response detectable in transcripts but not yet scored.
Section 115 Trust depletion. The council shifted from contributing $2M/year to the pension pre-funding trust to withdrawing $3–6M/year to balance an operating budget the City Manager calls structurally unsustainable. An A+ member has objected to this trajectory on the record. Gap: vote-based; not yet wired into scoring.
Each condition follows the same pattern: significant recurring expenditure or structural obligation; measurable standard the city's own documents establish; council response that falls short; member who does not challenge the gap. Absence of objection is not neutrality — it is ratification.
Council work is tiered by urgency and alignment with Berkeley's documented structural problems.
| Tier | Label | Definition | Examples |
|---|---|---|---|
| P1 | Crisis work | Directly addresses a documented structural problem (fiscal deficit, infrastructure backlog, pension liability, core services) | Budget reprioritization, fiscal referrals, no votes on spending, demand for efficiency data |
| P2 | Beneficial delivery | Legitimate city function with a clear, measurable objective — not in crisis, but real | Housing project approvals, specific public art commissions, parks maintenance, public safety equipment |
| P3 | Discretionary / ceremonial | Within Berkeley's general authority but low-priority, unaccountable, or purely performative | Cultural festivals, proclamations, arts grant programs, out-of-jurisdiction resolutions |
The Priority 1 (P1) engagement test: A member who generates only P2 activity — approves contracts, shows up for votes, stays out of trouble — but never engages with the P1 crises documented in the city's own reports is not doing the full job. A member who generates P3 activity while P1 problems go unaddressed is actively substituting low-priority work for the high-priority work the city's own documents say is urgent.
Scoring implication: Members with zero P1 referrals authored and low fiscal engagement (fiscal vote presence + fiscal concern rate) receive a low P1 engagement penalty of up to −0.10 applied to the composite grade. The penalty scales with the engagement gap — a member with no P1 referrals but high fiscal concern rhetoric and consistent vote attendance receives no penalty; a member with no referrals, low rhetoric, and frequent absences receives the full penalty.
Consent calendar items — passed en bloc without floor debate — are classified into five tiers that feed into the agenda scoring and Priority 1/2/3 (P1/P2/P3) analysis.
| Class | Label | What it means | Examples |
|---|---|---|---|
| 1 | P1 core | Directly addresses a documented structural crisis | Police staffing, fire equipment, infrastructure contracts, budget amendments |
| 2 | P2 delivery | Legitimate city function, clear deliverable | Housing project approvals, specific public art commissions, park improvements, fleet maintenance |
| 3 | P3 discretionary | Within city authority but low priority or performative | Cultural festivals, arts grant programs, proclamations, council office budget relinquishments |
| 8 | Administrative necessity | Has to happen regardless; minimal policy content | Minutes approval, bid solicitations, routine contract renewals, CalPERS side letters, grant applications for ongoing programs |
| 9 | Questionable scope | City doing what another body should do, or shouldn't be doing at all | County health service duplication, out-of-jurisdiction resolutions, non-competitive personnel rules, vague consulting without deliverables |
Art commissions: Specific public artwork contracts for named artists at named locations are class 2 (city delivers public art). Arts consulting, strategic planning, and grant-award programs are class 3 (program overhead). Grant acceptance for ongoing programs is class 8 (administrative).
Class 9 vs. class 3: Class 3 items are low-priority but within Berkeley's appropriate scope. Class 9 items represent scope the city arguably should not carry — county service duplication, political advocacy outside city authority, and governance structures that insulate dysfunction from accountability.
Consent calendar classifications are stored in agendas/classified/consent_items_classified.csv (594 items, Dec 2024–Apr 2026). Distribution: 1=41, 2=268, 3=58, 8=174, 9=53.
Action calendar classifications are stored in agendas/classified/action_items.csv (175 items, Dec 2024–Apr 2026). Distribution: 1=40, 2=47, 3=17, 8=56, 9=15.
Classification is topic-tier, not quality-of-response. A Priority 1 (P1) topic handled badly (e.g., a ballot-measures funding discussion that defaults to new taxes without contemplating cuts) is still class 1 — it addresses the right problem. The pipeline scores the quality of the response separately via fiscal concern rhetoric, new revenue preference signals, and vote record. Classifying it as P3 would undercount the council's P1 engagement; the penalty for the wrong answer belongs in the scoring layer, not the classification layer.
Three plain-language facts that explain the grade. No jargon.
check_fiscal_understatement() in agenda_scraper.py; scored −0.015/item authored, −0.007/item cosponsored, capped at −0.04. The failure may reflect naivete (not understanding what the proposal will require), weak effort (not thinking it through), or disregard for transparency — the penalty applies regardless of motive because the result is the same: the council and public lack a realistic cost picture when the item is heard.Metrics that require some familiarity with how city government works, but are explainable in a sentence.
Character = 0.35×collegiality + 0.25×humility + 0.20×warmth + 0.20×(1 − self-referential appeals)Constituency Preference Gap = 0.40×non-core speech% + 0.30×self-referential appeals + 0.30×(1 − fiscal engagement)The primary implemented example of special interest alignment scoring.
- Source: Transcripts + agenda cosponsorship
- Scale: 0 (reform-oriented) → 100 (status-quo aligned)
- What it measures: How invested is the member in the existing homeless services apparatus — $21.7M+/yr across 33+ programs, Housing First mandate, low-barrier ideology — versus demanding accountability, outcomes data, and reform?
- Why it matters: HSA is the primary implemented instance of the general outcomes-accountability principle (see A+ Ceiling Conditions). Berkeley's homeless spending has grown for a decade with no measurable reduction in visible homelessness. The same standard — demand metrics, question the model, require accountability for results — applies to any major program, but this is where the evidence record and transcript signal are richest. A member who champions more spending and resists outcome metrics is not representing the taxpayers who fund the program.
- Score distribution: See current scorecard output. High scores indicate deep alignment with the existing homeless services apparatus and resistance to accountability reform; low scores indicate a reform or outcomes-focused orientation.
Detailed signals for readers who want to understand the methodology or propose changes.
(absences/total)^1.5 × 0.25. This is lenient for 1–2 absences (−0.013 and −0.038) and severe for 4–5 absences (−0.120 and −0.151). The design reflects the judgment that one missed vote can have a real excuse; missing most major votes signals a different disposition. Maximum attendance deduction: −0.25.City Auditor reports are a separate evidentiary stream from incidents. They are not scored directly — the council's response to an audit is what gets scored.
Why this distinction matters: Incidents capture specific moments of individual behavior. Auditor reports establish independent, documented ground truth: facts the council is obligated to know, findings that create a duty to act, and a record that removes the ability to claim ignorance. When a council member subsequently acts in a way that contradicts a documented audit finding, that action is now interpretable as a choice — not a gap in knowledge.
How audits feed into scoring:
1. An audit is released and enters the registry (audit_findings.json) with key findings, auditor recommendations, and the action a taxpayer-aligned council should take.
2. When the audit goes on the council agenda, the council's first response is recorded. "Receive and file" without a companion motion is the baseline failure; a substantive motion by any member is the positive signal.
3. Subsequent council actions that contradict audit findings are scored harder because the ground truth is documented. A vote for a fifth bond cycle after receiving an audit that identified GF underfunding as the root cause of street decay is a different act than the same vote without that record.
4. Incidents that are grounded in an audit cite the registry key via the audit_ref field in incidents.json.
Ordinary voters will not read City Auditor reports. The audit findings registry (scores/pdfs/audit_findings.pdf) does that work — it translates what each audit found, what was warranted, and what the council actually did into a readable record. The pattern across audits — findings received, filed, and converted into bond campaigns — is the accountability story the scorecard is designed to tell.
Audits currently tracked (see audit_findings.json):
- streets_rocky_road_2025 — Rocky Road streets audit (Oct 2025): response documented; council received and filed, then directed a $300M bond with no GF reprioritization motion
- homeless_response_team_2025 — HRT audit (Jul 2025): pending council action; findings primarily staff-operational
- financial_condition_2026 — Financial condition audit (Apr 2026): pending council action; $32–33M structural deficit, 66% pension funded ratio, $1.8B unfunded capital, GFOA policy gap
Transcripts and agenda records capture what happens on the dais and in formal meetings. They do not capture everything that matters.
Incidents are documented accountability events that cannot be measured by transcript keyword analysis or voting records alone. An incident can arise from any verifiable source: meeting transcripts, constituent communications, news coverage, court filings, public records requests, auditor reports, or other public documents. Incidents are not limited to behavior inside formal proceedings — a newspaper article, a court filing, or a public records violation can all be the basis for an incident.
Incidents are scored on named dimensions, each of which carries a pillar tag indicating which aspect of a council member's performance is implicated:
A single incident can implicate multiple pillars: a conflict-of-interest failure affects Character & Conduct; if that conflict also caused public funds to flow to a contractor without proper oversight, it affects Fiscal Stewardship as well. Each dimension is scored and pillar-tagged independently, and its score rolls into its tagged pillar at full weight — there is no splitting or dilution across pillars.
Incidents can be positive as well as negative. A member who proactively surfaces a conflict of interest, demands a performance review, or votes no on a contractor reauthorization pending an audit earns points here.
Members with no incidents on record are not assumed clean — it means nothing has been documented yet.
A separate, automated incident system (incidents.json) tracks behavioral patterns observable in transcripts and agenda records:
Incidents are stored in incidents.json with structured fields:
| Field | Contents |
|---|---|
category |
One of seven categories (see below) |
date |
Date or approximate period |
description |
Plain-language description of the behavior |
source |
How the behavior was observed or established |
scoring_impact |
Suggested adjustment to composite score (typically −0.10 to +0.10) |
evidence_tier |
A / B / C — see below |
audit_ref |
Optional. Key into the auditor registry when the incident is a response to a specific audit finding |
The strength of a claim should be proportional to the strength of the evidence behind it. Each incident is assigned a tier that determines its effective weight in scoring:
| Tier | Definition | Weight |
|---|---|---|
| A | Primary public record: agenda items, vote records, official city emails, member official statements and farewell letters | 1.00 |
| B | Reputable reporting or personal newsletters: Berkeleyside, Berkeley Scanner, member personal-domain email | 0.75 |
| C | Direct observation or author knowledge: scorecard author's firsthand account without contemporaneous documentation | 0.50 |
Incidents with an audit_ref receive an additional 0.50× multiplier on top of their tier weight. The audit mechanism already penalizes the council's failure to act on a finding; the incident captures the specific member's behavior within that context. Applying full incident weight on top of the audit penalty would double-charge the same underlying failure.
| Category | Direction | What it captures |
|---|---|---|
revenue_without_cuts |
Negative | Sought new revenue without first asking what can be cut or done more efficiently |
performative_engagement |
Negative | Symbolic engagement — held meetings or sought input after decisions were made; performative not deliberative |
alternatives_dismissed |
Negative | Explicitly closed off alternatives without analysis or evidence |
claimed_ignorance |
Negative | Claimed not to know something they were obligated to know |
union_deference |
Negative | Sided with city unions without requesting productivity data or efficiency tradeoffs |
fiscal_integrity |
Positive | Pushed back on spending, demanded cost data, or advocated for cuts |
constituent_service |
Positive | Genuinely responsive constituent engagement with demonstrated follow-through |
Editorial incidents (the publicly documented incident pages) carry dimension-level pillar tags. Each dimension's score rolls into its tagged pillar at full weight. A dimension tagged Character & Conduct affects the Character pillar; a dimension tagged Fiscal Stewardship affects the Fiscal Stewardship pillar. If a single incident has dimensions in two pillars, both pillars are affected independently — there is no dilution. A per-pillar cap of ±0.30 prevents any single incident from zeroing out any one pillar.
Behavioral incidents (incidents.json) feed into the automated composite via the Fiscal Stewardship Alignment component. Each incident's raw scoring_impact is multiplied by its tier weight before being summed. Incidents linked to a City Auditor finding via audit_ref receive an additional 0.50× multiplier: the audit mechanism already penalizes the council's failure to act on a finding; the incident captures the member's specific behavior within that context.
Weighted behavioral incident totals are capped at ±0.30 before being applied to Fiscal Stewardship Alignment. The cap prevents any single member's incident record from dominating the overall score.
Members with no incidents in the log are not assumed clean — it means nothing has been documented yet. A member with many incidents has a richer evidentiary record.
The ±0.30 cap is correct for ordinary incidents. It prevents a heavily documented member from being punished merely because more evidence was collected about them, and it limits the effect of any single editorial judgment.
It has one consequence worth stating plainly: once a member reaches the cap, additional incidents have no effect, and severity becomes invisible. Twelve accumulated minor lapses and one catastrophic failure both land on the same floor. A member at −0.30 from a decade of small things scores identically to a member at −0.30 from a single event that misrepresented a public-safety finding to the Council.
That is the wrong result. It is addressed not by removing the cap but by removing a specific class of event from the capped pool entirely — see Material Governance Failures below.
Most incidents are additive: they accumulate, and the cap keeps that accumulation proportionate. Some failures are not additive. They are threshold failures — single events that call a member's governing judgment, honesty, or fitness into question in a way that no amount of routine good conduct offsets.
A Material Governance Failure is scored as a discrete event outside the ordinary incident pool. It bypasses the ±0.30 cap, cannot be canceled by unrelated positive incidents, is attributed individually by role, and may impose a grade ceiling as well as a direct composite deduction.
This extends a mechanism the scorecard already uses. A+ Ceiling Conditions (above) establish that some conduct cannot be fully compensated for by strength elsewhere. Material Governance Failures apply the same logic further down the scale.
An event enters this lane only when all four conditions hold. Any one failing returns it to the ordinary pool.
| Condition | Requirement |
|---|---|
| High-stakes action | Public safety, emergency response or evacuation, substantial expenditure, taxation or debt, legal exposure, conflict of interest, material contractor or program oversight, or a difficult-to-reverse commitment |
| Material premise failure | A factual, professional, legal, or evidentiary premise offered in support is unsupported by its cited source, outside the scope of the cited review, contradicted by an authoritative source, stated with materially greater certainty than the evidence supports, materially incomplete, or presented as completed analysis that has not occurred. The premise must matter to the decision — a peripheral error does not qualify. The audience may be the Council or the electorate — see below. |
| Authoritative notice | Before final action, the member receives notice sufficient to understand the premise may be wrong: a department head's statement, City Attorney opinion, auditor finding, staff report, written correction, or a direct statement in open session |
| Practical opportunity to cure | Before final action, the member could have corrected the document, amended the motion, disclosed the uncertainty, postponed, referred the matter, voted no, or supported a substitute preserving analysis |
A premise offered to the electorate in support of a ballot measure is treated the same as one offered to the Council, and the duty is if anything higher. A councilmember who signs a ballot argument is not commenting; they are asking voters to authorize taxation or debt on the strength of what the argument says.
Qualifying material includes signed ballot arguments and rebuttals in the official voter information guide, official campaign statements, and formal City communications about a measure. These are Tier A: they are published, attributed, and archived.
The test is the same — was the statement supported, complete, and within the scope of what the evidence established? — and so are the limits. A measure's supporters are entitled to argue for it, to emphasize its strengths, and to be wrong about predictions. Persuasion is not misrepresentation. What does not qualify as ordinary advocacy is characterizing the choice before voters in a way the measure's own terms do not support.
Why this is scored at all. A decision made on a false premise can be revisited when the premise is corrected. A decision authorised by voters on a false premise cannot: the mandate persists, and it is subsequently cited as evidence of what the public wanted. Misleading the electorate does not merely produce one bad outcome — it launders every downstream failure as the public's own choice. Where a governance failure is later defended by reference to voter approval, the accuracy of what voters were told is part of the record.
Seven factors, each scored 0, 1, or 2. Raw severity is their sum, 0–14.
| Factor | 0 | 1 | 2 |
|---|---|---|---|
| Decision centrality | Peripheral | Materially supported the action | The stated or indispensable basis |
| Stakes | Low | Meaningful policy or financial consequence | Safety, major expenditure, or legal exposure |
| Evidentiary clarity | Genuinely disputed | Strong contrary evidence | Direct contradiction from an authoritative primary source |
| Notice before action | None meaningful | Member had reason to know | Explicit correction before the decision |
| Opportunity to cure | None practical | Limited | Clear opportunity to correct, amend, postpone, or vote no |
| Failure to cure | Corrected before action | Partially addressed | Claim stood and action proceeded |
| Durability | Reversible at low cost | Creates a multi-year obligation | Creates a long-lived or practically irreversible obligation |
Why durability is scored separately. A Strategy for Strategy identifies the recurring Berkeley failure as commitments that outlive the assumptions and the money that justified them: grant deadlines substituting for evidence packages, pilots made permanent without their promised evaluation, capital built without an operating plan. A decision that is wrong and reversible is a mistake; a decision that is wrong and locks the City in for a decade is a different kind of failure. Folding irreversibility into "stakes" made the two score identically. It is now its own factor, and a maximum-severity finding requires it.
| Raw severity | Event classification |
|---|---|
| 0–4 | Routine — ordinary incident pool |
| 5–7 | Serious — ordinary incident pool |
| 8–11 | Major Governance Failure |
| 12–14 | Critical Governance Failure |
Raw severity describes the event. It does not yet describe any member's responsibility.
Responsibility is assigned by role. Nine members participating in one event receive one event record and nine attribution records — never nine duplicate incidents.
| Role | Multiplier | Definition |
|---|---|---|
| Originator | 1.00 | Authored, introduced, or formally asserted the premise |
| Post-notice defender | 0.85 | Repeated or defended the premise after correction |
| Co-sponsor | 0.80 | Lent their name to the item itself |
| Post-notice ratifier | 0.50 | Voted to proceed after correction without seeking a cure |
| Attempted cure, then ratified | 0.35 | Supported a motion that would have cured the failure, then ratified when it failed |
| Pre-notice supporter | 0.25 | Supported before correction, took no position after |
| Passive nonparticipant | 0.00 | Absent or recused |
| Corrective actor | 0.00 | Sought correction, delay, or analysis; or voted against proceeding |
effective_severity = raw_severity × role_multiplier, then classified on the same 8.0 / 12.0 thresholds. A member's role can therefore produce a different classification from the event's.
| Individual classification | Composite adjustment | Grade ceiling |
|---|---|---|
| Serious participation | Ordinary pool only | None |
| Major Governance Failure | −0.05 | B+ |
| Critical Governance Failure | −0.10 | C+ |
| Proven deliberate deception, or second unmitigated critical failure | −0.15 | D+ |
Why both a deduction and a ceiling. The deduction measures how much the event should move the composite. The ceiling prevents unrelated strengths — attendance, collegiality, routine fiscal signals — from restoring a top grade after a threshold failure. Without it, a member could retain an A-range grade following a critical integrity failure, which would misstate what the grade means.
When a ceiling lowers the displayed grade below the calculated composite, both are shown. The numerical score is never silently altered to match the ceiling; the normative judgment is made explicit and auditable.
A false representation is not a lie. Proven deliberate deception requires clear evidence that the member knew the representation was false or materially unsupported, intended reliance on it, and repeated or concealed it despite that knowledge — contemporaneous emails, prior written acknowledgments, admissions, or deliberate alteration of source language.
Absent that evidence, findings use language such as material misrepresentation, unsupported representation, claim contradicted by the responsible official, or material premise left uncorrected. The scorecard does not state that a member lied merely because a representation was false.
A Material Governance Failure replaces ordinary incident scoring for the same conduct. The same act is not scored twice under two labels. Independent signals measuring genuinely separate acts — the underlying vote record, attendance, a later refusal to correct the record, concealment or retaliation after the fact — may remain.
Unrelated positive conduct cannot offset a Material Governance Failure. Good constituent service, attendance, and sponsorship of useful unrelated measures do not purchase forgiveness for a threshold failure. Only conduct responsive to the same failure mitigates it:
| Remediation | Effect |
|---|---|
| Cured before final action | Returns to ordinary incident treatment |
| Prompt correction after action, plus meaningful reconsideration | Deduction halved; ceiling raised one band |
| Partial correction or disclosed uncertainty | Deduction reduced 25% |
| Admission without corrective action | Deduction reduced no more than 25%; ceiling stands |
| Later analysis happens to validate the decision | No mitigation |
| Defensive restatement, blame-shifting, or silent deletion | No mitigation |
Outcome does not erase process failure. A subsequent favorable result does not retroactively validate a decision made on a misrepresented evidentiary basis. That a design later passes review does not establish that it had been reviewed when the Council was told it had.
Material Governance Failures do not use ordinary incident decay. They hold full weight for the term in which they occurred:
| Age | Weight |
|---|---|
| 0–2 years | 1.00 |
| >2–4 years | 0.75 |
| >4–6 years | 0.50 |
| >6 years | 0.25 |
Documented remediation can reduce a failure faster than time does. Passive passage of time is not evidence of reform.
A Major or Critical finding must rest primarily on Tier A evidence — agenda records, official recommendations, meeting transcripts, department-head statements, staff reports, auditor reports, vote records, or public records. Tier B reporting may supply context but cannot independently establish the finding. Tier C observation is never sufficient.
Every Major or Critical finding must publish: the material claim; the source cited for it; what that source actually established; the authoritative contradiction; when each member received notice; what cures were available; what each member did afterward; the severity scoring factor by factor; each member's role multiplier; and any remediation considered.
A single event can produce a finding against most of the council at once. That is not a malfunction of the method, and it should not be read as one.
Berkeley's council votes as a block on 92.5% of recorded vote events. A body that acts as one unit fails as one unit. When a failure is genuinely council-wide — an item nobody amended, an opportunity nobody took — a council-wide finding is the accurate description of what happened.
These records are not a stack ranking. Individual attribution exists to distinguish roles within a shared failure — who authored it, who co-sponsored, who ratified, who tried to cure it — not to sort members against one another. A member with no attribution on a given event was absent, recused, or acted correctively; it does not mean they governed better in general, and a member carrying a ceiling is not thereby the worst member on the council.
The scorecard's per-member grades answer a different question from these event records. Where the two are in tension, the event record is the more specific claim and carries the citations.
A Material Governance Failure is not published unless a public explainer page accompanies it. This is enforced in the build: publish.sh refuses to deploy when an event's explainer file is missing.
The reason is that the score is not the finding. A grade ceiling a reader cannot trace to a sourced narrative is an assertion, not accountability. Each explainer is written for someone with no prior knowledge of the subject and must contain a dated chronology; the material claim with its cited source and what that source actually established, quoted verbatim; the authoritative contradiction or the omitted fact; what each member did and when; a full source list; and an explicit section stating what the finding does not establish, disclaiming any motive not in evidence.
That last requirement does the most work. A reader who can see the limits of a claim can judge the claim.
This classification requires Tier A evidence, satisfaction of all four qualification conditions, factor-by-factor scoring, individual role attribution, a published explanation, an explicit anti-double-counting review, and consistent application regardless of whether the scorecard author agrees with the policy outcome.
It must never be applied because the author disagrees with the result. The trigger is the failure of the governing process under consequential conditions — not the merits of the decision that process produced.
Events are recorded in governance_failures.json, one event object with per-member attribution records.
A separate penalty applies to members who were present at a formal audit presentation, voted to receive and file, and produced no follow-up motion within the response window. This is distinct from an incident: it is an automated signal for undifferentiated silence. Members who have a documented incident with an audit_ref matching the audit are exempt — their behavior post-audit is already individually characterized, whether positive or negative. Members in the follow_up_authored_by list in audit_findings.json are also exempt.
The silence penalty is −0.04 per audit event, applied to Fiscal Stewardship Alignment before the composite calculation.
The full incident catalogue is rendered as a separate PDF (scores/pdfs/incidents.pdf) for sharing with readers who want to see the underlying evidence behind scores.
These facts are not per-member scores but inform what "taxpayer-aligned" means in Berkeley's specific context. Any member who does not publicly challenge these structural problems is implicitly endorsing them.
rhetoric_action_gap_score pipeline variable partially captures this gap but could be refined.FISCAL_REFERRAL_VOTES (a curated list of upstream steps toward bond/tax ballot measures) and scored via score_fiscal_referral_votes(): −0.03 per item authored, −0.01 per item supported as cosponsor or aye vote, capped at −0.09. Weighted into taxpayer alignment as of April 2026.There is a further point that does not depend on any document. Candidates seek this office knowing what the office entails. A Berkeley councilmember is elected to govern a city carrying a nine-figure unfunded pension liability, a billion-dollar infrastructure backlog, and retiree health plans funded as low as 6.16%. Not knowing that is not a defense; it is a failure to prepare for a job actively sought.
The consequence for scoring is specific and deliberately narrowing. For conduct after March 2021, the question is rarely whether a member was on notice about the operating baseline, the pension trajectory, or the infrastructure backlog — they were, on a recurring cycle of the Council's own making, and they stood for election on that basis. What remains at issue is whether the record put before the body on a particular item was complete and accurate. That is a separate test, it is the one that does the real filtering, and it is where a finding against an individual member has to be argued.
- Vacancy savings were the balancing mechanism, and they are spent. Berkeley balanced recent budgets substantially by holding positions vacant and booking the salary savings. The FY2026 adopted budget assumed an 8% vacancy savings rate; a May 2025 analysis identified 104 vacant General Fund positions worth roughly $19.9M, of which about $9.3M was captured by holding them vacant. The City's Finance Director characterized this on the record on 20 May 2025: "these are just short-term fixes. Our unfunded liabilities for pension is about over $600 million, so $26 million is dwarfed." Two consequences follow, and both matter for reading any spending item. First, a vacant position is not a saving — it is a deferred cost, because the position returns to the baseline unless it is actually eliminated. Second, authorising new permanent positions during this period is worse than its headline cost implies: it enlarges the recurring baseline that the vacancy cushion was covering, at the point where the cushion is exhausted. By FY2027-28 the city was eliminating rather than merely freezing — 138 proposed position reductions, 100 vacant and 38 filled. This is context for scoring, not itself a per-member finding.
- No member spends meaningful time on the documented structural problems. The p1_speech_pct signal measures the share of a member's speaking time in turns containing the specific vocabulary of a documented structural failure — "structural deficit," "CalPERS," "pavement condition index," and similar. It is deliberately tight: a member saying "budget" or "roads" generically does not register. Measured across the corpus, seven of nine members sit at or below 1.0%, four are at or below 0.5%, and the council-wide maximum is 4.9%. This is the answer to a question often asked the other way round. The instinct is to fault individual members for time spent on matters outside Berkeley's authority — and off-mission speech does run 17% to 31%, which is high. But off-mission share does not separate members much, and the member most often criticized for it sits fourth of nine. What separates nobody, because nobody does it, is engagement with the problems the City's own auditor has documented. The finding is council-wide and is recorded as such rather than scored against individuals.
- This council votes together, so votes alone cannot tell members apart. Measured across 307 recorded votes: 86.3% were unanimous, median pairwise agreement between members is 95.7%, and the least-aligned pair on the council (Blackaby and Lunaparra) still agrees 90.7% of the time. Only 42 votes in the period were contested at all. This is the single most important fact about reading the scorecard: a member cannot be characterized by what they voted for, because almost everyone voted for almost everything. The differentiating question is how each member arrived at the shared vote — what they asked first, what they proposed instead, whether they moved a substitute, whether they held a stated position when it came under pressure. Each member's scorecard therefore opens with an at-a-glance line answering that question, and the supporting detail follows. Note also that dissent is not by itself a virtue: Lunaparra is on the losing side of contested votes more than twice as often as any other member and sits second-lowest on taxpayer alignment. Direction matters, not frequency.
- Per-member characterizations are checked against the other eight before being written. A claim that sounds damning often turns out to be the council norm. Three characterizations were cut during drafting for exactly this reason: "reaches for new revenue more than others" was drafted for Taplin, who is in fact second to Ishii; "unfocused" was drafted for Bartlett, whose 29% off-mission speech is matched by two members and exceeded by a third; and "the council's most verbose member" was inherited from an earlier summary and is false — the spread in average turn length runs 33.5 to 20.2 words and the member in question sits third. What survives comparison is what gets published.
- The record is incomplete by design, and gaps are marked. Waiting for a complete record would mean publishing nothing: the corpus is 61 transcripts and 68 annotated agendas against a council that has met for years, committee participation is largely unrecorded, and authorship metadata stops in March 2026. Every member scorecard therefore carries an explicit notice that absence of a finding means the finding has not been documented, not that none exists, together with an invitation to submit corrections. Silence in the record is never presented as a verdict.
- The Crayon Box — a standing test, applied by hand. Berkeley's council paints with two or three crayons when five or six are in the box. Closing a structural deficit admits at least six approaches: new taxes and new bonds, which are in constant use, and reprioritization, efficiency review, service-level reduction and provider comparison, which are rarely lifted. On any item that proposes new revenue, the standing questions are: which alternatives were available; which were raised; by whom; and if none were raised, is that recorded? Which crayons a member reaches for is a readable signal of disposition. More diagnostic is inconsistency: a member who applies efficiency review to one program and reaches for new revenue on a structurally identical one may be applying no consistent principle. That is not decisive on its own — it is a bad smell that warrants investigation, and it is logged as a question, never scored as a finding.
- The Crayon Box — competing motions are the evidence. The reliable instrument is not what members say but what they choose when a genuine alternative is on the floor. A substitute motion forces every member to pick, on the same item, in the same minute, on the record. That controls for everything the speech measures could not: a roll call has one name per vote, so staff presentations cannot contaminate it; the motions are known quantities, so polarity is unambiguous; and one vote is one datum, so repeated discussion cannot inflate it. Three distinct findings come out of it. (1) Unanimous silence. No member raises any alternative — that is blindness, whether intentional or not, and it is a council-wide condition rather than an individual failure. (2) Revealed preference. An alternative was offered and lost; the roll call shows exactly who reached for it. The Hopkins item is the worked example — an analysis-first substitute lost 4-5, a second substitute removing the unreviewed alignment lost 3-6, and the main motion carried 7-2. (3) The switch. A member states a principle and then votes against it on the same item. This is the strongest of the three, because the member supplied the standard themselves, so the finding requires no external benchmark and cannot be answered with "you simply disagree with them." Findings of type (3) are recorded against the member's own stated position, quoted verbatim, with the vote that contradicted it. (4) Mislabeling the crayon. The first three concern which option is chosen; this one concerns a false claim about what an option is. Telling the council that a configuration carried Fire Department support when the Fire Chief has stated on the record that it was never reviewed is not a narrow palette — it is handing the body a blue crayon labeled green. This is the one that crosses into Material Governance Failure territory, because it corrupts the record every other member is choosing from, which is exactly what qualification condition 2 (material premise failure) exists to catch. It also explains the role multipliers: the originator applies the false label, and a post-notice ratifier keeps drawing after being told the color is wrong. Narrow selection is a disposition that can be argued about on the merits; mislabeling is not.
- The Crayon Box — why it is not automated (negative result, Aug 2026). Keyword detection for the four unused options was built (CRAYON_BOX_KW in council_scorecard.py) and tested against all 55 transcripts. It does not work and is deliberately left unwired. Over the raw corpus the patterns mostly match staff, not members — the budget-presentation feed says "we're proposing to eliminate that position" dozens of times. Restricted to attributed member speech the contamination persists, because attribution folds long staff presentations into the chair's turn: the Budget & Finance chair scored 27 hits against a revenue baseline of 1, of which 22 are staff-voice constructions delivered by the City Manager's team. Polarity is also unhandled — "reallocating the funding into investing in our seniors" is a reprioritization proposal in the defund direction, and crediting it as fiscal discipline would mislead. And repetition inflates counts: one referral discussed across a single item produced six of nine provider-comparison hits. The root cause is shared with two abandoned governance-failure keyword screens: keyword lists test prose, while the question is whether an option was put on the table — an act, not a vocabulary. The lists are retained as a search tool for finding passages to read by hand, which is how the published finding was assembled.
- Substitute motions are not captured in structured data. The competing-motion test above must currently be built by hand from transcripts. Of 58 annotated agenda files, 17 record a vote line and none records substitute-motion text — the annotated agenda captures the final disposition, not the alternatives that were offered and defeated. The defeated substitute is precisely the evidence the test needs. Every finding of this kind therefore requires reading the item turn by turn, which is how the Hopkins record was assembled. A parser that extracts substitute motions and their roll calls from transcripts would make the test systematic rather than artisanal, and is the single highest-value addition to the pipeline.
- Speaker attribution: staff presentations no longer count as the chair's speech (fixed Aug 2026). Berkeley's transcripts label shared-microphone turns by room ("Boardroom:", "Redwood Room:") rather than by speaker, so attribution runs a state machine that holds the current speaker across turns. That is right when the floor passes between councilmembers and wrong the moment it passes to staff: the call-on detector only resolves council names, so a chair handing off to the City Manager left the floor pinned on the chair and every sentence of the ensuing budget presentation landed in that member's corpus. Measured by staff-voice constructions per 10,000 attributed words, the distortion tracked who chairs rather than what anyone said — Kesarwani 5.79, Taplin 2.96, Ishii 2.65, against Humbert 0.38. Three guards now apply: turns in unmistakable presentation voice are dropped from the member's corpus, an explicit hand-off to a named staff title releases the floor, and "Cypress Room" was added to the room labels (523 turns had been discarded entirely). After the fix the same measure reads Kesarwani 2.11, Taplin 0.45, Ishii 0.20, Bartlett 0.19, Humbert 0.00. Kesarwani's residual is genuine: she presents her own budget supplementals, where "we are proposing" means her co-sponsors. The hand-off release is deliberately bounded to six turns. An unbounded version was worse than the original bug — a chair running public comment produces no call-on pattern ("Next is Wendy A.", "your time's up"), so the floor never returned and 783 of Blackaby's own turns were discarded, an 18.7% cut to his corpus. Any future tightening should be tested against a chair running public comment, not only against a budget presentation. Effect on published scores was small and uniform: every composite moved by less than 0.005, with no reordering and no grade changes. Net corpus size rose 1.0%, because recovering Cypress Room added more than the guards removed.
- Material Governance Failure — historical backfill: The two-lane model was adopted in August 2026 and first applied to a July 2026 event. A full pass over the incident record was completed on 2026-08-04, driven by reader annotation of all 45 incidents rather than by keyword search — the earlier keyword screens were abandoned as unsound, once for using a severity threshold derived from the very measure under test. Five events now qualify; five are logged as screened and rejected in governance_failures.json -> screening_log, each with the condition that failed. The absence of a finding for an earlier event should still be read as "not assessed" unless it appears in the screening log.
- Notice must precede the act. Condition 3 is where most candidate failures die, and the chronology is checked before anything else. An authoritative finding released after the conduct it describes cannot supply notice for that conduct, however precisely it characterizes it — the 2026-04-08 financial-condition audit against the Oct–Dec 2025 bond activity is the reference case. Two consequences follow. First, genuinely bad decisions taken in good ignorance are not Material Governance Failures, and are recorded as Ownership of Outcome or as ordinary incidents instead. Second, an authoritative report sitting uncalendared is not yet notice to anyone: non-response to a report the City Manager has not scheduled is a scheduling fact, not a council decision. Such reports are held as watch items under pending_investigations with the qualifying standard written down in advance of the meeting, so the test is set before the outcome is known.
- Ownership of Outcome — sampling bias in the "shown" ledger: Seven of nine members currently carry zero documented instances of ownership shown, and the two exceptions earned theirs at a single meeting examined turn by turn. The ledger is built by investigating failures, so engagement that occurs in ordinary business goes unrecorded, and the dimension currently describes where the author has looked closely as much as how members behave. Bayesian smoothing toward a neutral prior keeps a thin record from reading as a damning zero, and members below three instances are not scored at all — but the top score being 50% is an artefact of coverage, not a finding about the council. Reading the per-instance lists is more informative than comparing the percentages until several more meetings have been examined at the same depth.
This document is the authoritative description of the scoring methodology. If you want to propose a new signal, adjusted weight, or different framing:
Send proposals to the maintainer. The underlying pipeline (pipeline.py) is modular — adding a new scoring function is straightforward once the signal is well-defined.