NextFin News - British police built a sprawling predictive-policing system around Avon and Somerset Police’s data science program, but the most important finding in the underlying records is not that the effort was ambitious. It is that the force could not convincingly show that several of its risk models were trustworthy enough to support the claims attached to them. The program used 13 risk models between 2017 and 2024, including tools tied to missing people, antisocial behaviour, and assessments of who might commit or fall victim to crime.
That matters because predictive policing is easy to expand and hard to verify. Once a score starts shaping attention, prioritization, or referrals, the model is no longer a technical prototype. It becomes an operational instrument. In Avon and Somerset’s case, the records show a system that grew across multiple use cases while key performance data and review notes raised doubts about transparency and reliability. Some of those concerns were serious enough that at least two risk-scoring models were later abandoned after Bristol City Council staff said they could no longer trust them.
The issue also lands at a sensitive moment for UK policing. The government has created PoliceAI, a £75 million-backed body hosted by the College of Policing to help roll out AI tools across 43 police forces in England and Wales. The direction of travel is toward more automated policing, not less. That makes the Avon and Somerset case less a local oddity than a stress test for the standards that should apply before a police force turns models into practice.
Predictive systems in policing are often sold as a way to move from reaction to prevention. In theory, that means better resource allocation and earlier intervention. In practice, it requires more than a dataset and a score. It requires proof that the model performs better than crude alternatives, that it is stable over time, and that its errors are understood well enough to avoid amplifying existing biases or mistakes. The public records reviewed in the source material point to exactly that gap: a widening operational footprint, but a thinner-than-needed validation story.
A Sprawling System Was Harder to Trust Than To Build
The first lesson from the Avon and Somerset program is that scope is not the same as quality. The force’s data science effort produced a large collection of scores and tools, not one neatly bounded application. The available performance data included more than 36,000 model scores disclosed to the source material, and those records were reviewed by an independent analyst and an AI auditing firm. That scale can create the illusion of rigor. It can also make failures harder to see, because no single score tells the whole story.
That is where predictive policing often goes wrong. A model can appear useful if it seems to align with what officers already believe or with where police have historically concentrated attention. But if the underlying patterns are just policing history folded back into a new interface, the model is not discovering risk so much as relabeling old patterns. The source material indicates that some of Avon and Somerset’s outputs were judged to show genuinely poor predictive performance, which is a stronger critique than merely saying the system was controversial. It means the outputs themselves failed a basic test.
The fact that at least two models were later abandoned after Bristol City Council staff said they could no longer trust them reinforces the same point. Once a public body stops trusting a score, the burden shifts to the system’s owners to explain what went wrong: the data quality, the model design, the evaluation method, or the governance around its use. If that explanation is absent, the apparent sophistication of the tool becomes beside the point.
This is why the case is so instructive for readers who think about policing as a public-service function rather than a software deployment. A police model can be lawfully assembled from legal gateways and administrative data, but legality does not answer whether the resulting system is accurate, proportionate, or legitimate. Reviewers in government records captured that distinction directly: “Legality is not the same as legitimacy.”
“Legality is not the same as legitimacy,” reviewers said in government records obtained in the underlying records reviewed for the story.
The phrase is blunt because it captures the central tension. Police forces are tempted to equate access to data with the right to use it operationally. But data access is only the beginning. Before a score affects a human decision, the force needs to show that the score works, not just that it can be produced. The Avon and Somerset materials suggest that the program reached the second point without fully clearing the first.
Why Predictive Policing So Often Outruns Its Evidence
The second lesson is more structural: predictive policing grows faster than its evidence because its benefits are easier to imagine than to prove. Police leaders facing stretched budgets and public pressure want tools that promise to anticipate harm. Vendors and internal data teams can usually produce enough plausible-looking output to justify continued experimentation. But the validation standard required for real operational trust is much higher than the standard needed to keep a pilot alive.
That gap matters because forecasting in policing is not a laboratory exercise. Crime patterns shift. Records are incomplete. Staff use tools differently from one team to another. A model that looks acceptable in a limited evaluation can degrade once it is used across different neighbourhoods, incident types, or time periods. The more categories a system covers, the more opportunities there are for hidden failure. Avon and Somerset’s program spread across missing people, antisocial behaviour, and crime-risk assessments, which means it was not dealing with one stable prediction problem but several different ones at once.
That is also why the number of models matters. Thirteen separate risk models between 2017 and 2024 is not just a count; it is a sign of operational sprawl. When a system grows by accumulation, accountability fragments. One team owns missing-person alerts, another owns antisocial behaviour scoring, another owns crime-victimization risk, and the public is left with the impression that “the algorithm” works even if some of its pieces do not. The source material’s performance records suggest that the system did not earn that confidence.
There is a further problem: models built from police-held information can inherit the history of where police already looked. That creates a circular dynamic. More patrols produce more records, more records make certain areas look riskier, and higher risk scores justify more patrols. Without careful controls, the model can end up amplifying what police already did rather than revealing what will happen next. That is one reason an apparently data-rich system can still underperform.
The source material indicates that more than 36,000 model performance scores were available for review, which sounds like a formidable evidence base. Yet quantity is not the same as comparability. A large dataset can still fail if the scoring method, target definition, or evaluation window is wrong. The important question is not how many scores exist, but whether they prove the system is better than common-sense alternatives. The review described in the source material suggests the answer, in several cases, was no.
PoliceAI Raises the Stakes for Governance, Not Just Technology
The third lesson is the most important policy-wise: the UK is moving toward more AI in policing even as the trust problem is still unresolved. PoliceAI, the £75 million-backed body hosted by the College of Policing, is designed to help roll out AI tools across 43 forces in England and Wales. That is a national scaling effort. It means the lessons from Avon and Somerset are not local housekeeping; they are a preview of the governance challenge ahead.
The danger is that national momentum can turn a disputed tool into an accepted one before the evidence matures. Once a technology is wrapped in the language of modernization, it becomes harder for police leaders to step back and ask whether the underlying use case deserves to survive. That is especially true when the tool promises efficiency. Faster triage, earlier intervention, and more targeted deployment are compelling goals. But if the score is unreliable, efficiency just means the force will move faster in the wrong direction.
The case also highlights a familiar asymmetry in public-sector AI. The people building and approving the system often know more about it than the people affected by it. That imbalance can be tolerated only if the system is clearly accurate and auditable. When the outputs are disputed, the legitimacy deficit widens quickly. In that sense, Avon and Somerset’s program is not just a story about data science; it is a story about institutional trust.
The quoted reaction from the source material is telling. The concern is not merely that a model is technical or opaque. It is that a model may be given power over people’s lives before it has earned that power. That is a strong standard, but policing is a strong-power institution. If a score influences where officers look, who gets contacted, or whose risk gets elevated, the public should expect proof of reliability that is stronger than the fact that the software exists.
“I don’t think an AI model should have that kind of power over people’s lives,” Pegram said.
That objection should not be read as anti-data. It is an argument for evidence before authority. Police forces can and should use data to understand risk, but the burden is on them to show that a model genuinely improves judgment rather than dressing up old assumptions in predictive language. The Avon and Somerset records suggest that burden was not met across the full program.
That matters well beyond one force because the UK is choosing scale before it has solved governance. If the new national push produces more systems like this one, the question will not be whether police are using AI. They already are. The question will be whether the public can trust the scores enough to accept the consequences.
The most revealing line in the story may be the simplest: a system can be large, sophisticated, and data-heavy, and still fail the one test that matters most. In policing, the real metric is not whether a model can make a prediction. It is whether anyone can trust the prediction enough to let it shape the world.
Explore more exclusive insights at nextfin.ai.

