Predicting which records match¶
In the previous tutorial, we built and estimated a linkage model.
In this tutorial, we will load the estimated model and use it to make predictions of which pairwise record comparisons match.
from splink import Linker, DuckDBAPI, splink_datasets
db_api = DuckDBAPI()
df = splink_datasets.fake_1000
df_sdf = db_api.register(df)
df_sdf.as_duckdbpyrelation().limit(5).show()
┌───────────┬────────────┬─────────┬────────────┬─────────┬─────────────────────────┬─────────┐
│ unique_id │ first_name │ surname │ dob │ city │ email │ cluster │
│ int64 │ varchar │ varchar │ date │ varchar │ varchar │ int64 │
├───────────┼────────────┼─────────┼────────────┼─────────┼─────────────────────────┼─────────┤
│ 0 │ Robert │ Alan │ 1971-06-24 │ NULL │ robert255@smith.net │ 0 │
│ 1 │ Robert │ Allen │ 1971-05-24 │ NULL │ roberta25@smith.net │ 0 │
│ 2 │ Rob │ Allen │ 1971-06-24 │ London │ roberta25@smith.net │ 0 │
│ 3 │ Robert │ Alen │ 1971-06-24 │ Lonon │ NULL │ 0 │
│ 4 │ Grace │ NULL │ 1997-04-26 │ Hull │ grace.kelly52@jones.com │ 1 │
└───────────┴────────────┴─────────┴────────────┴─────────┴─────────────────────────┴─────────┘
Load estimated model from previous tutorial¶
import urllib.request
from pathlib import Path
def get_settings_text() -> str:
# assumes cwd is repo root
local_path = Path.cwd() / "docs" / "demos" / "demo_settings" / "saved_model_from_demo.json"
if local_path.exists():
return local_path.read_text()
# fallback location for settings - the file as it is on master, for e.g. colab use
# TODO: update ref
url = "https://raw.githubusercontent.com/moj-analytical-services/splink/master/docs/demos/demo_settings/saved_model_from_demo.json"
with urllib.request.urlopen(url) as u:
return u.read().decode()
import json
settings = json.loads(get_settings_text())
linker = Linker(df_sdf, settings)
Predicting match weights using the trained model¶
We use linker.inference.predict() to run the model.
Under the hood this will:
-
Generate all pairwise record comparisons that match at least one of the
blocking_rules_to_generate_predictions -
Use the rules specified in the
Comparisonsto evaluate the similarity of the input data -
Use the estimated match weights, applying term frequency adjustments where requested to produce the final
match_weightandmatch_probabilityscores
Optionally, a threshold_match_probability or threshold_match_weight can be provided, which will drop any row where the predicted score is below the threshold.
df_predictions = linker.inference.predict(threshold_match_probability=0.1)
df_predictions.as_duckdbpyrelation().limit(5).show(max_width=10000)
Blocking time: 0.01 seconds
Predict time (post-blocking): 0.10 seconds
┌─────────────────────┬────────────────────┬─────────────┬─────────────┬──────────────┬──────────────┬──────────────────┬───────────────────────┬───────────────────────┬───────────────────┬──────────────────────┬───────────┬───────────┬───────────────┬──────────────────────┬──────────────────────┬─────────────────────┬──────────────────────┬────────────┬────────────┬───────────┬────────────────────┬────────────┬────────────┬────────────┬─────────────────────┬─────────────────────┬────────────────────┬─────────────────────┬──────────────────────────┬──────────────────────────┬─────────────┬───────────────────────┬───────────────────────┬───────────────────┬─────────────────┬───────────┐
│ match_weight │ match_probability │ unique_id_l │ unique_id_r │ first_name_l │ first_name_r │ gamma_first_name │ tf_first_name_l │ tf_first_name_r │ mw_first_name │ mw_tf_adj_first_name │ surname_l │ surname_r │ gamma_surname │ tf_surname_l │ tf_surname_r │ mw_surname │ mw_tf_adj_surname │ dob_l │ dob_r │ gamma_dob │ mw_dob │ city_l │ city_r │ gamma_city │ tf_city_l │ tf_city_r │ mw_city │ mw_tf_adj_city │ email_l │ email_r │ gamma_email │ tf_email_l │ tf_email_r │ mw_email │ mw_tf_adj_email │ match_key │
│ double │ double │ int64 │ int64 │ varchar │ varchar │ int32 │ double │ double │ double │ double │ varchar │ varchar │ int32 │ double │ double │ double │ double │ date │ date │ int32 │ double │ varchar │ varchar │ int32 │ double │ double │ double │ double │ varchar │ varchar │ int32 │ double │ double │ double │ double │ varchar │
├─────────────────────┼────────────────────┼─────────────┼─────────────┼──────────────┼──────────────┼──────────────────┼───────────────────────┼───────────────────────┼───────────────────┼──────────────────────┼───────────┼───────────┼───────────────┼──────────────────────┼──────────────────────┼─────────────────────┼──────────────────────┼────────────┼────────────┼───────────┼────────────────────┼────────────┼────────────┼────────────┼─────────────────────┼─────────────────────┼────────────────────┼─────────────────────┼──────────────────────────┼──────────────────────────┼─────────────┼───────────────────────┼───────────────────────┼───────────────────┼─────────────────┼───────────┤
│ 21.506834253252443 │ 0.9999996644187925 │ 28 │ 30 │ Thomas │ Thomas │ 4 │ 0.006016847172081829 │ 0.006016847172081829 │ 6.309947619380791 │ 0.040393298600093885 │ Gabriel │ Gabriel │ 4 │ 0.004884004884004884 │ 0.004884004884004884 │ 6.580128192319569 │ -0.10616211501540729 │ 1976-09-15 │ 1976-09-15 │ 2 │ 7.715603236685051 │ London │ London │ 1 │ 0.21279212792127922 │ 0.21279212792127922 │ 2.3410443182772083 │ -0.9401495213654414 │ gabriel.t54@nichols.info │ gabriel.t54@nlchois.info │ 3 │ 0.0025348542458808617 │ 0.0012674271229404308 │ 7.952135813932513 │ 0.0 │ 0 │
│ -2.7980062727528967 │ 0.125710472856419 │ 37 │ 860 │ Theodore │ Theodore │ 4 │ 0.012033694344163659 │ 0.012033694344163659 │ 6.309947619380791 │ -0.9596067013999061 │ Morris │ Marshall │ 0 │ 0.004884004884004884 │ 0.007326007326007326 │ -2.1597988656731415 │ 0.0 │ 1978-08-19 │ 1972-07-25 │ 0 │ -1.116036665159834 │ Birmingham │ Birmingham │ 1 │ 0.04920049200492005 │ 0.04920049200492005 │ 2.3410443182772083 │ 1.1725506113839206 │ t.m39@brooks-sawyer.com │ NULL │ -1 │ 0.0063371356147021544 │ NULL │ 0.0 │ 0.0 │ 0 │
│ -2.7980062727528967 │ 0.125710472856419 │ 39 │ 860 │ Theodore │ Theodore │ 4 │ 0.012033694344163659 │ 0.012033694344163659 │ 6.309947619380791 │ -0.9596067013999061 │ Morris │ Marshall │ 0 │ 0.004884004884004884 │ 0.007326007326007326 │ -2.1597988656731415 │ 0.0 │ 1978-08-19 │ 1972-07-25 │ 0 │ -1.116036665159834 │ Birmingham │ Birmingham │ 1 │ 0.04920049200492005 │ 0.04920049200492005 │ 2.3410443182772083 │ 1.1725506113839206 │ t.m39@brooks-sawyer.com │ NULL │ -1 │ 0.0063371356147021544 │ NULL │ 0.0 │ 0.0 │ 0 │
│ -2.7980062727528967 │ 0.125710472856419 │ 43 │ 860 │ Theodore │ Theodore │ 4 │ 0.012033694344163659 │ 0.012033694344163659 │ 6.309947619380791 │ -0.9596067013999061 │ Morris │ Marshall │ 0 │ 0.004884004884004884 │ 0.007326007326007326 │ -2.1597988656731415 │ 0.0 │ 1978-08-19 │ 1972-07-25 │ 0 │ -1.116036665159834 │ Birmingham │ Birmingham │ 1 │ 0.04920049200492005 │ 0.04920049200492005 │ 2.3410443182772083 │ 1.1725506113839206 │ t.m39@brooks-sawyer.com │ NULL │ -1 │ 0.0063371356147021544 │ NULL │ 0.0 │ 0.0 │ 0 │
│ 6.523156368990757 │ 0.9892443210014298 │ 47 │ 49 │ Erin │ Erin │ 4 │ 0.0048134777376654635 │ 0.0048134777376654635 │ 6.309947619380791 │ 0.3623213934874556 │ Rogers │ NULL │ -1 │ 0.006105006105006105 │ NULL │ 0.0 │ 0.0 │ 2010-01-02 │ 2010-03-01 │ 0 │ -1.116036665159834 │ London │ London │ 1 │ 0.21279212792127922 │ 0.21279212792127922 │ 2.3410443182772083 │ -0.9401495213654414 │ e.rogers3@hopkins.org │ e.rogers3@honkips.org │ 3 │ 0.0025348542458808617 │ 0.0012674271229404308 │ 7.952135813932513 │ 0.0 │ 0 │
└─────────────────────┴────────────────────┴─────────────┴─────────────┴──────────────┴──────────────┴──────────────────┴───────────────────────┴───────────────────────┴───────────────────┴──────────────────────┴───────────┴───────────┴───────────────┴──────────────────────┴──────────────────────┴─────────────────────┴──────────────────────┴────────────┴────────────┴───────────┴────────────────────┴────────────┴────────────┴────────────┴─────────────────────┴─────────────────────┴────────────────────┴─────────────────────┴──────────────────────────┴──────────────────────────┴─────────────┴───────────────────────┴───────────────────────┴───────────────────┴─────────────────┴───────────┘
Clustering¶
The result of linker.inference.predict() is a list of pairwise record comparisons and their associated scores. For instance, if we have input records A, B, C and D, it could be represented conceptually as:
A -> B with score 0.9
B -> C with score 0.95
C -> D with score 0.1
D -> E with score 0.99
Often, an alternative representation of this result is more useful, where each row is an input record, and where records link, they are assigned to the same cluster.
With a score threshold of 0.5, the above data could be represented conceptually as:
ID, Cluster ID
A, 1
B, 1
C, 1
D, 2
E, 2
The algorithm that converts between the pairwise results and the clusters is called connected components, and it is included in Splink. You can use it as follows:
clusters = linker.clustering.cluster_pairwise_predictions_at_threshold(
df_predictions, threshold_match_probability=0.2
)
clusters.as_duckdbpyrelation().sort("cluster").limit(10).show(max_width=10000)
Completed iteration 1, num edges remaining to process: 380
Completed iteration 2, num edges remaining to process: 24
Completed iteration 3, num edges remaining to process: 0
┌────────────┬───────────┬────────────┬─────────┬────────────┬────────────┬───────────────────────────┬─────────┐
│ cluster_id │ unique_id │ first_name │ surname │ dob │ city │ email │ cluster │
│ int64 │ int64 │ varchar │ varchar │ date │ varchar │ varchar │ int64 │
├────────────┼───────────┼────────────┼─────────┼────────────┼────────────┼───────────────────────────┼─────────┤
│ 0 │ 1 │ Robert │ Allen │ 1971-05-24 │ NULL │ roberta25@smith.net │ 0 │
│ 0 │ 3 │ Robert │ Alen │ 1971-06-24 │ Lonon │ NULL │ 0 │
│ 0 │ 0 │ Robert │ Alan │ 1971-06-24 │ NULL │ robert255@smith.net │ 0 │
│ 0 │ 2 │ Rob │ Allen │ 1971-06-24 │ London │ roberta25@smith.net │ 0 │
│ 4 │ 4 │ Grace │ NULL │ 1997-04-26 │ Hull │ grace.kelly52@jones.com │ 1 │
│ 5 │ 5 │ Grace │ Kelly │ 1991-04-26 │ NULL │ grace.kelly52@jones.com │ 1 │
│ 6 │ 6 │ Logan │ pMurphy │ 1973-08-01 │ NULL │ NULL │ 2 │
│ 7 │ 8 │ NULL │ Dean │ 2015-03-03 │ NULL │ NULL │ 3 │
│ 7 │ 10 │ NULL │ Dean │ 2015-03-03 │ Portsmouth │ evied56@harris-bailey.net │ 3 │
│ 7 │ 9 │ Evie │ Dean │ 2015-03-03 │ Pootsmruth │ evihd56@earris-bailey.net │ 3 │
└────────────┴───────────┴────────────┴─────────┴────────────┴────────────┴───────────────────────────┴─────────┘
10 rows 8 columns
The estimated cluster id is the cluster_id column, the true cluster id (which we only know because this data is labelled) is the cluster column.
Note that the value in the column isn't meaninful, all that matters is that it creates the right grouping. For instance, the three Evie Dean records are all assigned to estimated cluster_id 7, correctly grouping them together. It does not matter that the value used to group them together in the cluster column (the true label) is 3.
We can see the model has done a reasonable but imperfect job of estimating the true cluster. This is a simple model trained on a small and very messy dataset, so it is not surprising that accuracy is not better..
Further Reading
For more on the prediction tools in Splink, please refer to the Prediction API documentation.
Next steps¶
Now we have made predictions with a model, we can move on to visualising it to understand how it is working.