The situation
A county planner reviewing a variance application needs to know the complete permit and enforcement history for a parcel. That history exists: it was generated by the county's own permitting and code enforcement departments over decades. The problem is where it lives.
In most county offices, permit records from the 1980s through the early 2010s are scanned PDF files, organized by folder structure on a shared network drive, accessible only by staff who know the filing conventions. A single parcel inquiry may require opening five or more individual documents, each requiring the reviewer to find and read the relevant field.
This is a research task that has to happen in the middle of a more complex review. It pulls planners away from the substantive analysis they are trained to do.
Workflow before Apaluma
A typical parcel history lookup, before structured extraction, follows this sequence:
- Open the shared drive folder for the parcel's address or APN
- Identify which PDFs contain permits versus violations versus inspection reports
- Open each PDF and read through it to find the relevant fields
- Copy key information by hand into the review memo or GIS attribute table
- Repeat for each document type in the parcel's folder
- Cross-check against the current parcel database for address or APN discrepancies
For a parcel with a long history, this process takes 20 to 45 minutes. For a department processing several dozen variance or permit applications per month, this is a significant share of staff time allocated to a task that should take two minutes.
How Apaluma helps
Apaluma ingests the archive of scanned PDFs and extracts structured records from each document type. Each extracted record is matched to an APN in the county parcel layer and written to a structured data file alongside the parcel geometry.
The result is a GIS layer that can be opened in ArcGIS or QGIS alongside the base parcel layer. A planner clicking on a parcel sees its full permit and violation history as attributes, not as a folder of documents to manually review.
The extraction process includes a confidence score for each field. Records where the pipeline is uncertain about field values are flagged for staff review rather than silently included at the wrong confidence level. This means the layer can be trusted: flagged records are known unknowns, not silent errors.
What you get
At the end of an extraction run, the county receives:
- A GeoJSON, Shapefile, or KML file of extracted records linked to parcels
- A CSV summary with one row per record, suitable for spreadsheet review
- A flagged-record report listing records the pipeline was uncertain about, with the source page cited
- A field coverage summary showing extraction completeness by document type
These outputs are designed to load into the county's existing GIS environment with the same import steps as any other external data layer. No new GIS platform is required. No consultant is required for the initial import.
Starting a pilot
The pilot process is scoped before it starts. We ask for a sample of documents from one document type in your archive and confirm what the extraction pipeline can produce against your specific documents. This takes about a week and requires no ongoing access to your systems.
After the pilot, you receive the sample output and an accuracy report. If the results meet your needs, we discuss the full-archive scope and subscription. If they do not, we document the gap and you keep the sample output at no cost.
Contact Us to Start the Conversation