From scanned archive to GIS layer in four steps.
Apaluma handles the full conversion pipeline. No manual data entry, no GIS specialist required to get the first layer live.
Request Early AccessEvery step of the pipeline, explained
The pipeline runs automatically once you provide a document set. Here is what happens at each stage.
Upload scanned PDFs individually or as a batch zip. Apaluma accepts mixed document types in a single upload and queues them automatically.
Layout-aware OCR reads each page. The extraction engine identifies document type, parses key fields, handles handwritten annotations, and normalizes values to a consistent schema.
Each extracted record matches to an Assessor Parcel Number. Ambiguous or low-confidence matches are flagged for staff review with the original scan attached.
Matched records export as GeoJSON, Shapefile, KML, or CSV. Load the layer directly into ArcGIS, QGIS, or your web map platform without additional conversion.
What happens to uncertain records
Apaluma does not silently fill in gaps. Records that do not meet the match threshold route to a review queue with the source scan, so staff make the final call on ambiguous cases.
Each extracted field carries a confidence score. Fields below the threshold are flagged visually in the review interface so staff know exactly which values to verify against the original scan.
The review interface shows the extracted data fields alongside the original scanned page. Reviewers can correct individual fields without re-processing the full document.
Every parcel match records the method used, the confidence level, and any manual overrides. The audit trail is exportable for records management compliance.
Correct a parcel ID in your county database and re-run matching against the existing extracted records. No need to re-upload or re-process the original scans.
Questions about the process
Processing speed depends on document complexity and batch size. A county archive of several thousand permit PDFs typically completes its first extraction pass within a few business days. We scope processing time estimates during the pilot evaluation based on your specific document set.
The extraction engine handles typical government archive scan quality, including grayscale at 200 DPI or better, moderate ink bleed, and partial page shadows from physical binding. Documents significantly below 150 DPI or with extreme degradation will route to the manual review queue rather than producing unreliable extractions.
No GIS-side configuration is required. The exported layer loads directly into ArcGIS or QGIS using standard file import. We provide a setup call to walk your GIS staff through the first import and show how the parcel linkage fields align with your existing data model.
When the APN field is absent or illegible, Apaluma attempts address-based matching against your county parcel layer. If address matching also falls below the confidence threshold, the record is placed in the review queue with the source document. Your staff makes the final parcel assignment, which is recorded in the audit trail.
See the pipeline run on your document type
We scope each pilot around a specific document type from your archive. Apply and we will confirm what Apaluma can extract from your county's records.
Start a Pilot Evaluation