Heads up for priority items for running and developing the website for launch (as opposed to general backend iterative work to keep improving the systems). There are longer timeline tasks to support the database functioning and its role as a queryable resource for researchers at database/SCHEMA_BACKLOG.md. For the full set of improvement work across the backend, see the Roadmaps section of the root README, which gathers the deferred and in-progress roadmap kept across section READMEs.
While the website functions as is, there are some improvements that need to be made. These are, in brief:
- Rectify Community Water System (CWS) data so the CWS tier outcomes are fully supported. This means:
- understanding the final list of cws demand units
- designing and building the CWS schema to add the water-system layer and link it to the existing urban demand-unit tables (a draft design with open questions lives in the database README Planned tables section)
- adding the demand units the CWS team has given us that are still missing from the tables, both the
tier_locationids with no entity row and the rows that still lack geometry - reconciling the urban gw/sw flags
- reconciling the M&I delivery crosswalk
- rebuilding and republishing the Mapbox tileset via MTS so those polygons render on the map, because the map draws from the tileset and not the database, so a DB-only fix wouldn't reach the map. The MTS workflow is documented in the frontend map package README under "Adding new tile data via Mapbox Tiling Service"
- Get NOD/SOD locations correctly specified in the database and served by the API. Today that membership lives in frontend code instead of coming from the API, due to a quick fix to include NOD/SOD data categories in the radar chart for the last AC meeting.
- Harden the statistics variable lists and calculations. Source the lists from database tables instead of seed tables and lists in scripts, and harden the calculations that read them in collaboration with WAM team.
- Move demand-unit geometry off the entity rows into dedicated geometry tables, the way every other spatial layer already stores it, and stop the monthly audit export from dumping full geometry inline.
Goal: Update drinking-water/CWS delivery data so the CWS_DEL tier outcome is fully supported. The reference spreadsheets are staged in data/reference/cws/, though check with the CWS team to see if they have drifted since May 2026. Four sub-tasks remain:
- CWS dataset tables: Finalize and build the CWS schema. A draft design exists in
database/README.md, Planned tables (proposingcws_entity,cws_du_link, and a list registry), but it has open questions to resolve first, notably whether the CWS list registry should be new tables (cws_list/cws_list_du_member) or a reuse of the existingdu_urban_groupmechanism, and whether to normalize the utility name / PWSID out ofdu_urban_entity.community_agencyfree text into its own lookup (SCHEMA_BACKLOG § 8), since the draftcws_entityis one row per PWSID. A third is where any CWS geometry (system points or centroids) lives: design it into a dedicated geometry table, not onto thecws_entityrow, following the pattern in A5. To be clear, the CWS tables themselves do get built in the database as part of this task. What is deferred is only the separate migration that moves geometry off the existingdu_*_entitytables, since the map reads geometry from Mapbox tiles rather than the DB, so that migration is not on the CWS critical path. Resolve the open questions above, then create the CWS tables and load them from the staged spreadsheets. - Add the missing demand units, then refresh the map tiles: Confirm the final CWS demand-unit list with the team first (some entries may also need removing). Then three pieces of work:
- Add the ids that have no entity row at all. A handful of ids appear in the tier map (
tier_location) but have no entity row, so they have neither attributes nor a polygon: the seven CWS_DEL ids (ACFC,KCWA,MHILL_NU,SBCWD,SVWRD,TLMNE,UNION) plus the AG_REV id07S_PA. Those seven are the same set the database README calls the "master DUs missing from CalSim," listed there next to theESB355HHS-list entry as per-id accept-or-fix decisions (open questions). Decide accept vs fix for each, then add the rows. - Load polygons for the rows that have no geometry yet. The 72 ids that have a row but no polygon. This is the same coverage work as A4.
- Rebuild and republish the Mapbox demand-units tileset (
coeqwal.calsim_demand_unitsrecipe) via MTS once the geometry is loaded. The map draws from the tileset and not from the DBgeom, so a database-only fix never reaches the map. The MTS workflow is in the frontend map package README under "Adding new tile data via Mapbox Tiling Service".
- Add the ids that have no entity row at all. A handful of ids appear in the tier map (
- gw/sw value reconciliation: Each urban DU has two flags,
gwandsw, recording whether it draws groundwater, surface water, or both. The same flag pair appears in three places that do not fully agree: the committed seed CSV (du_urban_entity.csv, the values loaded in the DB today), the curated CalSim Table 3-7 rollup, and the M&I team xlsx. The reconcile script (verified working, it reproduces the May 2026 baseline) and the walkthrough lay the three side by side. It finds 32 seed-vs-xlsx disagreements. Five of them are plain seed errors, rows where both CalSim and the xlsx agree on a value the seed got wrong (03_PU1,24_NU4,26N_NU5,26N_PU1,26S_PU2); correct those five rows in the seed CSV now. The rest are rows where the seed already matches CalSim and only the xlsx differs, so they are a data-team tier-rule call, not a data fix. Two smaller follow-ups remain on top of that: sign-off on the03_PU3override and fillingNAPA2in the rollup. All of this is itemized in the walkthrough roadmap; the WAM and Drinking Water/CWS teams are the right reviewers for these decisions. Separately, urbangw/sware still stored asVARCHARwhile the ag and refuge tables are alreadyBOOLEAN. That type migration is independent of the value reconciliation (walkthrough Step 7, SCHEMA_BACKLOG § 6). - Master crosswalk vs
du_urban_variable: Reconcile the SW-DU M&I delivery-variable crosswalk againstdu_urban_variable(buckets, script status, per-step detail), then merge the updated delivery-variable crosswalk in.
Concepts: DU vs CWS vs contractor in water_user_categories.md.
Goal: NOD/SOD membership for demand units should live in the database and be served by the API. Today the website carries this classification in frontend code instead of reading it from the API.
Next step:
- Backfill the 5 empty
du_urban_grouprows (nod,sod,swp_served,cvp_served,swp_delivery_point). Urban is already wired:/demand-unitsjoins the membership table, defaults totier, and accepts?group=, so backfilling is enough to make?group=nod(etc.) return data. - Add the ag-side
du_agriculture_group/du_agriculture_group_membertables, mirroring the existingdu_urban_groupshape in the ERD (Layer 03). Membership is derivable fromcs3_type/provider/hydrologic_region_idondu_agriculture_entity. The ag endpoint has nogroupparameter yet, so this also means adding the?group=join to/ag-demand-units, not just the tables. - Do the same for refuge if it needs NOD/SOD slicing (no
du_refuge_grouppair orgroupparameter exists yet); otherwise scope this task to urban + ag. - Optionally, return group-membership flags on the
/demand-unitsand/ag-demand-unitsAPI rows so the website filters client-side in one fetch, then drop the frontend's hardcoded copy. This is separate from the server-side?group=filter.
Open decisions (full discussion in the database roadmap): whether to derive NOD/SOD membership from a rule (SAC -> NOD, SJR + TULARE -> SOD) rather than hand-curating rows, including a WAM-team ruling on the four regions that don't fall cleanly on one side of the Delta (DELTA, SOCAL, NC, EXPORT); and whether to consolidate the du_urban_group mechanism with the cws_list registry so we don't ship two ways to define membership. The reservoir endpoints are the worked example of the *_group pattern this task extends.
Goal: Source the statistics variable lists from database tables (Layer 04) instead of hardcoded Python dicts and seed-CSV prefix derivation, and harden the calculations that read them. Today only DU Urban reads its mapping from SQL (du_urban_variable). MI, CWS, Delta, the reservoir thresholds, and the AG aggregate / GW-only sets are Python constants.
Next step:
- Build out the Layer 04 variable tables (
reservoir_variable,mi_contractor_variable,cws_aggregate_variable,ag_variable,refuge_variable,delta_variable) and have each statistics module read from the DB (variable-list migration). - Work the calculation-review items (reservoir spill threshold and volume, and the WAM-team needs-review backlog).
Where we are: Some tier locations do not fully resolve to an entity row (attributes) and a polygon (geometry). CWS_DEL is the bulk of the gap and is handled in A1. The piece unique to A4 is the AG_REV side: 07S_PA has no du_agriculture_entity row.
Next step: Add the 07S_PA ag row (accept-vs-fix as with any gap). For the CWS_DEL gaps, follow the procedure in A1, the same add-rows, load-polygons, republish-tileset flow, tracked together with A1.
Home: geometry.md: the coverage scorecard (Coverage), the tier locations with no geometry (Referenced by tier_location but not present), and the full missing-id list (Appendix).
Now: DU geometry sits inline on the entity rows. du_urban_entity, du_agriculture_entity, and du_refuge_entity each carry a full geom (MultiPolygon) plus a redundant geom_wkt text mirror and srid. Only network_gis and reservoir keep geometry in a dedicated table (wba and compliance_station are inline too, but small and low priority). The DU footprints are the largest and their rows are the most heavily read for non-spatial data, so they are the ones to move. Nothing in the API or frontend reads the DB geometry today, the live map draws polygons from the Mapbox tileset, so it is latent at runtime. See how geometry is represented now.
Target: move DU geometry into a dedicated table keyed by (du_id, du_class), FK-enforced the way network_gis links to network, drop the geom_wkt mirror, and design CWS/PWS geometry the same way. See how geometry should be represented.
Next step:
- Move
geom/geom_wkt/sridoff the threedu_*_entitytables into a dedicated geometry table (the patternnetwork_gisandreservoiruse), then drop the columns. Migration sketch and pickup steps are ingeometry.mdand SCHEMA_BACKLOG § 6f. - Relocate the misplaced project rows while here: 11 rows whose
du_classis Agricultural or Refuge (10 ag, 1 refuge) sit indu_urban_entity. They got there because the backfill migration01g_add_missing_du_urban_entities.sqlmisguidedly added missing CalSim project DU (PU,PA,PR) todu_urban_entityfor completeness (they all have delivery arcs in CalSim), recording each row'sdu_classcorrectly but leaving the ag and refuge ones in the urban table. Each row self-identifies viadu_class, so move them todu_agriculture_entity/du_refuge_entity. While here, also resolve26N_NA, which is duplicated across the urban and ag tables.