Berkeley was warned not to measure too much, told how to connect programs to priorities, and shown how ranking could guide cuts. Thirteen years later, Council has rediscovered the first instruction and is still missing the rest.
Berkeley publishes performance measures, and has in every budget book since fiscal year 2022. The question is whether those measures show that a program is producing the result it exists to produce.
Mostly they do not. They record volume — inspections completed, permits processed, outreach contacts made, tickets closed, miles installed. Those are useful for managing workload. They answer what did we do, not what changed because we did it.
Activity measurement asks what government did. Outcome measurement asks what changed because government acted. A city can do more and more of something without the underlying condition improving at all — and it will not be able to tell.
| What gets counted | What would need to be known |
|---|---|
| Homeless outreach contacts made | Share of clients still housed after twelve months |
| Potholes repaired | Change in Pavement Condition Index |
| Miles of bicycle infrastructure installed | Change in serious cyclist injuries, adjusted for exposure |
| Building permits processed | Median end-to-end approval time and applicant cost |
| Trees planted | Canopy still standing after five years |
Prioritizing means ranking, and ranking means knowing what each program delivers for what it costs. A city relying primarily on activity data has no defensible way to distinguish a program that works from one that merely stays busy.
Faced with a roughly $29 million General Fund shortfall, the City Manager asked every department to model reductions at a common percentage. That method assumes every activity is equally important and equally able to absorb a cut, and the city’s own budget documents describe the consequences in specific departments as severe.
To protect one service and cut another deeper, someone has to be able to say, on the record, that the first delivers more per dollar than the second. Absent outcome data, nobody can say that. Without a defensible ranking system, flat percentage cuts become the administratively easiest option.
The same gap disables the rest of the toolkit. STOP — the decision to cease an activity altogether — requires evidence that it is not working. Without that evidence, stopping becomes much harder to justify and the cost base tends to persist. Exploring alternatives requires something to compare alternatives on. Deciding whether a purchase was worth making requires knowing what it has delivered since. This is the condition underneath them.
The City Auditor found this problem directly in one of Berkeley’s largest discretionary spending areas.
Berkeley has no outcome tracking system for its homeless programs. The city cannot assess whether they are working.
The programs in question account for more than $21.7 million a year, spread across some 33 programs. The Auditor recommended implementing outcome tracking, and requiring performance benchmarks before renewing program contracts.
City Auditor, Berkeley Homeless Response Team: Coordination, Oversight, and Outcome Tracking, July 2025. The Auditor is elected by Berkeley voters and reports independently of the City Manager and the Council.
That is one department’s programs, and it does not establish that every Berkeley program lacks outcome measures. It establishes that the problem is real and consequential in the largest discretionary spending area the city controls.
In February 2026 the Auditor followed with a research report, A Guide to Measuring Performance in the City of Berkeley, framed as opportunities for management consideration rather than a finding of non-compliance. Berkeley’s measures are scattered across documents, some data is hard to find, and peer cities use standardized reporting periods, public dashboards and in some cases outcome-based budgeting. Berkeley’s police report 911 response times in their annual report; the figure does not appear in the budget book.
The number exists. It is not where a decision-maker would look for it.
Stand on the platform at Downtown Berkeley BART and there are two numbers that matter: minutes to the Richmond train, minutes to the San Francisco train. Occasionally a third — delayed, cancelled.
Those numbers are useful precisely because of everything behind them. Under “5 min” sit trains, operators, schedules, signals, maintenance records, staffing rosters, control systems and live position data. When the number becomes 25, nobody at BART convenes a committee to invent the information needed to find out why. The arrival board is the front end of an operating system that already exists.
Berkeley is now proposing its own arrival board: ten to twenty measurable goals, updated quarterly, on a public dashboard.
The question is not whether the right number is ten, twenty or thirty. It is what sits underneath them.
Ten to twenty may well be the right number of gauges. Five would leave too much out; a hundred would guarantee nobody watches any. The 2013 report discussed below made the same point, warning that cities attempting priority-based budgeting commonly “try to measure too much” and that too much information “can lead to analysis paralysis among key decision makers.”
The referral solves a real dashboard problem, but not the underlying management architecture. The number answers how many indicators can we watch. It does not answer what is this a number of.
Suppose Council selects two: serious crime, and the condition of city parks. Crime worsens; park condition improves. Nothing follows from that alone. Council cannot set arrests against lawns mowed and determine where the next million belongs, and knowing that Police hit 85% of one target while Parks hit 105% of another does not establish which produced more public value. Heterogeneous raw metrics do not compare. The comparison has to run through agreed strategic objectives, service standards, marginal benefit, statutory obligation, equity and cost.
The top-level measure becomes useful only as the visible end of a hierarchy:
City strategic objective → desired outcome → contributing departments and programs → program objectives → output, outcome and efficiency measures → spending and staffing → trend against target → resource decision.
Council does not need to compare arrests with lawns. It needs to decide the relative weight of a handful of strategic outcomes, see which programs contribute to each, know what level of service each is producing, and know what changes when money is added or removed.
The operational test is drill-down. When a top-level indicator turns red, a councilmember should be able to click through to the contributing programs, their component measures, their costs and their trends — in the meeting, not six weeks later. Staff should not have to build a fresh presentation each time someone asks why a number moved, because department managers should already be using that information to run their departments.
If a dashboard shows homelessness outcomes deteriorating and the next step is “refer to staff to determine why,” the city has built an elegant way to discover that the underlying management information still does not exist.
In 2013 the City Auditor’s Office asked UC Berkeley’s Goldman School of Public Policy to examine alternatives to Berkeley’s modified across-the-board General Fund budgeting. The resulting student report, Berkeley Based Budgeting, was transmitted to Council by Auditor Ann-Marie Hogan in April 2014. It was not an audit finding; it was a detailed implementation blueprint Berkeley had asked for:
| Recommendation | What it asked Berkeley to build |
|---|---|
| 1. Establish a strategic vision | A short Statement of Community Priorities, built from surveys and forums rather than only from who attends meetings, with each priority defined as specific outcomes |
| 2. Reorganize functions around programs | An inventory of what the city actually does, organized so that every activity advances at least one stated priority |
| 3. Develop performance measures for each program | Measures tested for validity, reliability, responsiveness, ease of understanding, economy of collection and balance — and kept deliberately minimal |
| 4. Allocate funding by performance and priority | Programs scored against the priorities and ranked, so that reductions could be targeted rather than spread evenly |
The fourth recommendation described the quartile method: score programs against the community priorities, sort them, and cut by rank. It cited Grand Island, Nebraska, which needed roughly $1 million in savings and took more from the bottom quartile while shielding the top two — noting approvingly that the final budget departed from the target in places, because officials retained the judgment to protect a critical program that scored badly. Its stated benefit: it “avoids using the across-the-board cutting method on every program.”
Berkeley’s practice, the report said, was to ask each department for the same percentage reduction and then adjust the result “based on a process and criteria that do not appear in the Budget Documents, and are essentially invisible to the public” — the “black box” appearance of government. It predicted the position Berkeley now occupies: the existing system “has worked well for small or occasional cuts, but is ill equipped to address major structural changes.”
Thirteen years later the city faced a General Fund deficit of roughly $29 million in each of two years, asked departments for 10% and 12.5% reduction scenarios, and began a conversation about developing outcome measures that might someday inform allocation.
The FY2027–28 plan was not purely across-the-board: the budget document states that it “does not reflect across-the-board equal reductions” but a strategic approach, staff named criteria, and the Mayor ran a Council priorities exercise during the budget cycle. Berkeley has intent to prioritize. What it lacks is the operationalized machinery — the program inventory, the scored measures, the ranking — that would let a priority be applied to a dollar and defended afterward.
Berkeley never had to reject any of this.
| When | What arrived | How it was handled |
|---|---|---|
| April 2014 | Auditor-solicited priority-based budgeting recommendations, with costs, benefits and specified changes to the work plan and budget documents | Information Calendar — no vote required |
| April 2020 | Strategic Plan Performance Measures Pilot, built on “is anyone better off” | Information Calendar — received and filed |
| April 2026 | Referral for ten to twenty outcome goals, quarterly updates, public dashboard | Consent — zero speakers, zero letters |
Not one of the three produced a moment at which somebody had to say yes or no on the record. The Information Calendar requires no vote; consent requires no discussion. The 2020 pilot used Results-Based Accountability’s core question: is anyone better off? The annotated agenda records its disposition in three words:
“Received and filed.”
Annotated Agenda, April 14, 2020
The 2014 report was actionable as delivered, each recommendation carrying its own costs, benefits and implementation steps. No adoption, referral, work-plan change or follow-up item followed it.
A rejected proposal leaves a record, an argument and a reason. A proposal routed to a calendar that requires no vote leaves nothing, and can arrive again a decade later as though for the first time.
The 2026 referral carries the same defect forward: it asks for quarterly updates and a dashboard once goals exist, but sets no deadline for the goals themselves. Nothing is ever late, and no meeting arrives at which someone must account for the absence. Across seventy-four annotated agendas in this project’s record the phrase “measurable goals” appears on one — and in the three months and eight meetings after adoption, including the adoption of an entire biennial budget, no member asked about it once.
Performance measures did appear in the budget book from fiscal 2022, which is institutional progress and should be credited. But appearing in a budget book is not the same as governing by them, and the record does not show appropriations being made contingent on results in the years since.
The test is not whether a dashboard appears. It is whether the measures change what the Council does. For each major program, that means a stated public condition it is meant to improve, a baseline, a target, a date, the cost of reaching it, a reporting frequency — and a stated consequence if performance does not move. Build that, then put the ten or twenty most important gauges on the front page.
The April 2026 referral is incomplete in a specific way: it treats selection of the gauges as the substantive work, when the gauges become useful only after the underlying system exists. Berkeley has been told this before, by a report its own Auditor commissioned, and has had three opportunities to decide about it without ever being required to decide.
Berkeley does not need another collection of numbers demonstrating that city employees are busy. It needs a system capable of telling Council and residents what their government accomplished with the money it spent — and helping them decide what the next dollar should buy.