# SIG-P4B 2026-07-30 — Session Notes

Participants:
  - rafa (UTC+1)
  - PAtwater | UTC-8 | 31P

---

## 📖 The Reading
Not clearly identified in the discussion. This session was not a reading discussion but a working **interview**: rafa interviewed PAtwater to gather material for a report on water-data standardization (with the intent of feeding the transcript into Claude to help draft the report). The conversation references prior sessions covering case studies (GTFS, UK Open Banking, electricity data) and a "water rate discoverability" project associated with "Maxwell"/"Matthew."

## 🧭 Overview
rafa interviewed PAtwater to develop an anchoring story for a report on why water-rate data standardization is valuable, why it's painful, and why past standardization efforts stalled. PAtwater walked through his early career (hand-scraping ~200 utilities' water rates as an intern), the purpose of rate benchmarking, the failed/fizzled standardization initiatives in California, and the structural reasons standardization never fully took hold. The arc lands on a thesis: standardizing at the data-extraction level is a dead end, but new AI tooling (à la "Maxwell's" approach) could lower the cost of compliance enough to unlock the long-promised value.

## 💡 Key Points & Themes

**The origin story / anchor for the report**
- **PAtwater:** ~a decade in water data. First job was an internship at a large water utility hand-scraping the water rates of ~200 Southern California utilities over a summer — cheaper than paying a consultant (~$50–100k) for a periodic survey.
- The takeaway that stuck with him: "this is harder than it should be," even two years after Google existed.

**Why rate benchmarking matters**
- The goal was to construct a "typical bill" (e.g., 10,000 gallons/month) and see the range and trend of costs across Southern California — essentially benchmarking.
- As a wholesaler (drawing from the Sierra Nevadas and Colorado River), the utility wanted to understand the downstream impact of its prices on retail customers, alongside other cost inputs (power, construction, supply chain, chemicals).
- Water is relatively cheap (~$50–80/month plus sewer), so it's usually not a top-10 public issue — except during droughts or lead-in-pipe crises (Flint referenced). But boards still track it because customers and board members care about rising bills.

**How the data is actually used**
- Rate benchmarking feeds **board decisions and strategic/business planning** (1–5 year cycles): are rates rising too fast/slow, how do we compare to peers, what can the community tolerate given infrastructure and cost pressures.
- **rafa** reframed the core driver as strategic scenario planning (3% vs 5% vs 10% increases) more than individual customer complaints — PAtwater largely agreed. Rate hikes of ~6% vs 3% feel dramatic to the public even though they're not "2x or 10x."

**Simplifications and data limits**
- The benchmarking never captured full rate schedules — they defined "typical" usage and simplified (e.g., single-family only), because full rate structures (single/multifamily/commercial/industrial, water-budget rates) are too complex to present.
- Getting data was time-intensive: if not on the web, you called customer service, who often didn't have the full picture.

**Why cross-utility visibility is wanted (and resisted)**
- Every utility has local visibility but lacks clear comparison to neighbors; a natural board question during rate-setting is "how do our rates compare?"
- Price variation is driven by legitimate factors (hydrology, geography, pumping needs, expensive desalination — the San Diego/Escondido example) **and** by a spectrum of good-to-poor management. Sensitivity around the data stems partly from not wanting poor management exposed.

**Standardization efforts and why they fizzled**
- **California Water Data Consortium** (created ~5–7 years ago); a **water reporting project** started ~2022 aimed at harmonizing standards — produced a report and "fizzled."
- **Open Water Rate Specification** (~2017–2018): won a state open-data challenge, proved the concept, could flexibly represent full rate structures, was used in a Cal-Nevada water rate survey with several hundred utilities specified. But it never usurped existing alternatives; the GitHub repo is ~7 years out of date.
- "Maxwell" reused this specification (refactored slightly, now as JSON) on the back end.

**Barriers to adoption**
- Technical/capability friction (GitHub was "too far" for ~98% of water staff; four-to-five steps was too much), staff time cost, and plain inertia ("we've always reported it this way; it's what's legally required").
- A tragedy-of-the-commons dynamic: you only need ~40% adoption to do your own strategic planning, and you don't need to be in that 40%.
- State-level "IT inertia" — legacy "Frankenstack" systems (built ~10 years ago, changed only via expensive vendors like Deloitte) make it costly to change reporting.

**The current frontier / "local minimum"**
- Utilities are legally required to submit rate data (Urban Water Management Plans via Excel-into-web-portal; the **Electronic Annual Report** via web form dumping into a giant shared text file).
- The submitted data is comprehensive but **not fit for purpose**: it captures a "typical bill" input, not the full rate schedule, and is bad both from a data-management and a domain perspective.
- **The "original sin":** the system (dating to ~2006) asked for narrow inputs (the "10,000 gallons" figure), then accreted individually-reasonable questions over decades into an unusable, ossified form. rafa's analogy: they built a spreadsheet of hard-coded values instead of one with formulas.

**The AI / paradigm-shift thesis**
- If AI can lower the cost of specifying data (take a rate structure / PDF and output the standard spec), you get all the long-dreamed-of benefits without the manual work — "throw the work at the machine."
- **rafa's framing of the core problem:** the price of fixing standardization is higher than its perceived value, so it's always "priority #7 (or 10 or 15), not #1." The Google Maps analogy: standardization only got solved there because being absent from Maps became an obvious problem with direct value.
- The opportunity: 2026 ≠ 2006 — the cost of solving may now be much lower and the value much higher. But solving standardization *at the data-extraction level* is still the wrong solution.

## 🔀 Questions & Disagreements
- rafa repeatedly pushed a "thought experiment": would a simple yes/no compliance flag (e.g., rates within ±10% of CPI) suffice? PAtwater pushed back — building that boolean wouldn't be easier, water is political, and board members would argue endlessly over the threshold (why 10% vs 5% vs 15%). He emphasized the granular data point is what board members actually need.
- Some back-and-forth over **where the pain sits** — local utilities (repeatedly asked for duplicative data, staff time) vs. the state (keeps asking for more reports because it doesn't get usable data). PAtwater's answer: both.
- Minor unresolved factual points: why Buffalo, NY water is so expensive (~$42/thousand gallons) — PAtwater had "no idea"; Escondido/San Diego high cost hypothesized (not confirmed) as the desalination plant.
- Naming/attribution ambiguity in the transcript between "Maxwell" and "Matthew" for the AI-based rate project — the transcript uses both and it's unclear if they are the same person.

## 🔗 References Mentioned
- California Water Data Consortium (CDC/CADC)
- Open Water Rate Specification (~2018)
- California Water Data Challenge (state open-data challenge)
- Cal-Nevada water rate survey (~2017–2018 industry association survey)
- Urban Water Management Plans; Electronic Annual Report (state reporting mechanisms)
- "Making Conservation a California Way of Life" regulatory framework
- Case studies from prior sessions: GTFS (transit data), UK Open Banking, electricity data
- Google Maps / Yelp (analogy for value-driven data adoption)
- Flint, Michigan (lead-in-pipe crisis as example of when water becomes urgent)
- Deloitte (example state IT vendor)
- Pumped-storage hydropower (mentioned in opening small talk, not part of the report)

## ✅ Action Items & Next Time
- **rafa:** Feed this interview transcript into Claude to draft the report (the stated experiment for the session).
- **PAtwater:** Considering an audience/distribution list — former rate consultants, his former boss (a former CFO at "Met[ropolitan]"), and similar friendly readers who could "run with" the report.
- **rafa (idea floated):** Use the existing consultant "cottage industry" to advantage — run an awareness campaign asking utilities to demand a specific deliverable (LLM-readable rate info posted on their website) from consultants they already hire, potentially at little extra cost, and showcase one exemplary consultant report to raise the bar.
- No specific next reading or next meeting date was set in the transcript.

## ⭐ Memorable Quotes
- **PAtwater:** "From a global perspective, it's kinda sad. Right? Because you get worse information for more public money. But from a micro perspective..." (their sky won't fall from paying a consultant).
- **PAtwater (on AI):** "If you could just... put it into the standard spec, then you get all the benefits that we are just scheming and dreaming about, but you don't have all the work. You just throw the work at the machine."
- **rafa:** "The price is high and the value is ambiguous... this is always priority number seven or 10 or 15 on the list. It's not priority number one."
