America is a nation of local laws - millions of them. From noise ordinances and construction codes to rules about fishing licenses and roadside mango peddling, every town, city and county has its own set of regulations and unique way of doing things.
Accessing those laws can be a byzantine quest. Dense PDFs and outdated, inconsistent government websites can make it impossible to easily compare building codes or nuisance laws. For developers, it can make project costs soar. For residents, it can lead to confusion and frustration. And for researchers, it can stymie efforts to understand how laws are enforced or affect people.
A first-of-its-kind project at UC Berkeley is using AI and data science to help untangle this national patchwork by collecting every digitally accessible law in a free, open-access database. While online repositories have cataloged state and federal statutes for years, the new project is the first time such a database has existed specifically for local laws, from Alameda, California, to Zephyrhills, Florida. The existence of the new resource opens the door for others to generate custom databases and local chatbots that can quickly generate comparisons of local laws on topics spanning home construction to dog leash laws.
We've built a new kind of telescope - one pointed at governance itself.
Diag Davenport, UC Berkeley
"We've built a new kind of telescope - one pointed at governance itself," said Diag Davenport, an assistant professor of technology policy, governance and society at Berkeley and lead researcher on the project. "It could be a fundamental shift in the way people interact with the rules that govern their daily lives."
Davenport, who is part of the Goldman School of Public Policy and the School of Information, started thinking about the project six years ago while studying computer algorithms, systemic bias and the criminal legal system. Around that time, he wondered if he could invert that research focus and study bias in the actual law itself.
To answer that, he needed to study local laws around the country and use statistics to measure how they differed in places that had high levels of historic discrimination compared with those that did not. That's when he hit the first major problem: No database of all local laws existed.
The more he looked, the more he realized they were stored in a mishmash of places, from proprietary company databases to convoluted local government websites. Then there was the scale of the confusion: more than 3,000 counties and another 6,000 or so local jurisdictions that make their own laws.
"Of course, nothing's ever as simple as you want it to be," Davenport said. "Once you realize how fragmented it all is, it's easy to understand why no one's done the work."
The project simmered for years, but it took new life about a year ago when technological advances allowed the team to use one AI tool on top of another to do the heavy lifting.

Brandon Sánchez-Mejia/UC Berkeley
First, the researchers needed to collect municipal and county laws from thousands of government and third-party websites. Davenport and his colleague, Denis Peskoff, a postdoctoral scholar at Berkeley, consulted with lawyers as they developed the collection process and designed it to meet the technical requirements imposed by the sites hosting the documents. They also worked with Joe Barrow, an AI researcher specializing in documents and datasets, and Christopher Vu, a Berkeley undergraduate studying computer and data science.
But that was only the beginning. The larger technical challenge was turning nearly 10,000 often unwieldy documents - roughly 7 million pages, many stored as blurry, poorly structured or otherwise inaccessible PDFs - into data that researchers could actually use. The team used a vision-language optical character recognition model called LightOnOCR to extract both text and structure from the documents. Processing the archive required a massive amount of computing power, but the result was a relatively standardized body of machine-readable text that is far easier and less costly to search, analyze and use with modern AI systems than the original PDFs.
From there, the researchers used models from OpenAI to tag and organize samples of the laws and distilled that work into buckets that could be used across the full corpus.
The resulting database contains millions of unique ordinances from all 50 states. It allows researchers to search across local laws at a scale that was previously extremely difficult and to identify patterns in how communities regulate everything from housing and public space to business activity and everyday conduct.
"I think it was actually impossible to do this work until six or 12 months ago," Davenport said. "As far as I'm aware, we've done this at a larger scale than anyone else."

Brandon Sánchez-Mejia/UC Berkeley
The team published a version of the database online in June; some people have already created searchable interfaces based on the team's work.
And last week, they learned a paper describing their project was accepted for the Conference on Neural Information Processing Systems, one of the top conferences focused on machine learning and AI. They'll present their findings in December.
The next stage of the work is even more exciting, Davenport said. Computer scientists and other tech-savvy users can use the vast repository of information to train chatbots for specific subject areas. As an example, one could be trained on Bay Area building codes and quickly report the different requirements to build an apartment complex in Berkeley compared with El Cerrito. That could make it possible for developers to more easily build cost estimates into project plans or help a homeowner follow the rulebook for building a backyard accessory dwelling unit. Critically, Davenport said any forthcoming tools based on the data must be free and publicly accessible.
Journalists, researchers and policymakers will also have streamlined access to compare laws, saving countless hours that would otherwise be spent researching hard-to-navigate documents. Users can compare regional differences, study how difficult to understand laws are in some regions, and even probe regional differences in enforcement.
"Effectively, what we have are thousands of experiments about how to regulate housing, sidewalks and helmets," Davenport said. "The real promise here is making local government legible. Once people can actually see and compare the rules that shape everyday life, we can ask much better questions about which ones work, who they work for and what we might want to do differently."