Skip to content

Predicting which records match

Open In Colab

In the previous tutorial, we built and estimated a linkage model.

In this tutorial, we will load the estimated model and use it to make predictions of which pairwise record comparisons match.

from splink import Linker, DuckDBAPI, splink_datasets

db_api = DuckDBAPI()
df = splink_datasets.fake_1000
df_sdf = db_api.register(df)
df_sdf.as_duckdbpyrelation().limit(5).show()
┌───────────┬────────────┬─────────┬────────────┬─────────┬─────────────────────────┬─────────┐
│ unique_id │ first_name │ surname │    dob     │  city   │          email          │ cluster │
│   int64   │  varchar   │ varchar │    date    │ varchar │         varchar         │  int64  │
├───────────┼────────────┼─────────┼────────────┼─────────┼─────────────────────────┼─────────┤
│         0 │ Robert     │ Alan    │ 1971-06-24 │ NULL    │ robert255@smith.net     │       0 │
│         1 │ Robert     │ Allen   │ 1971-05-24 │ NULL    │ roberta25@smith.net     │       0 │
│         2 │ Rob        │ Allen   │ 1971-06-24 │ London  │ roberta25@smith.net     │       0 │
│         3 │ Robert     │ Alen    │ 1971-06-24 │ Lonon   │ NULL                    │       0 │
│         4 │ Grace      │ NULL    │ 1997-04-26 │ Hull    │ grace.kelly52@jones.com │       1 │
└───────────┴────────────┴─────────┴────────────┴─────────┴─────────────────────────┴─────────┘

Load estimated model from previous tutorial

import urllib.request
from pathlib import Path


def get_settings_text() -> str:
    # assumes cwd is repo root
    local_path = Path.cwd() / "docs" / "demos" / "demo_settings" / "saved_model_from_demo.json"

    if local_path.exists():
        return local_path.read_text()

    # fallback location for settings - the file as it is on master, for e.g. colab use
    # TODO: update ref
    url = "https://raw.githubusercontent.com/moj-analytical-services/splink/master/docs/demos/demo_settings/saved_model_from_demo.json"
    with urllib.request.urlopen(url) as u:
        return u.read().decode()
import json

settings = json.loads(get_settings_text())

linker = Linker(df_sdf, settings)

Predicting match weights using the trained model

We use linker.inference.predict() to run the model.

Under the hood this will:

  • Generate all pairwise record comparisons that match at least one of the blocking_rules_to_generate_predictions

  • Use the rules specified in the Comparisons to evaluate the similarity of the input data

  • Use the estimated match weights, applying term frequency adjustments where requested to produce the final match_weight and match_probability scores

Optionally, a threshold_match_probability or threshold_match_weight can be provided, which will drop any row where the predicted score is below the threshold.

df_predictions = linker.inference.predict(threshold_match_probability=0.1)
df_predictions.as_duckdbpyrelation().limit(5).show(max_width=10000)
Blocking time: 0.01 seconds


Predict time (post-blocking): 0.10 seconds


┌─────────────────────┬────────────────────┬─────────────┬─────────────┬──────────────┬──────────────┬──────────────────┬───────────────────────┬───────────────────────┬───────────────────┬──────────────────────┬───────────┬───────────┬───────────────┬──────────────────────┬──────────────────────┬─────────────────────┬──────────────────────┬────────────┬────────────┬───────────┬────────────────────┬────────────┬────────────┬────────────┬─────────────────────┬─────────────────────┬────────────────────┬─────────────────────┬──────────────────────────┬──────────────────────────┬─────────────┬───────────────────────┬───────────────────────┬───────────────────┬─────────────────┬───────────┐
│    match_weight     │ match_probability  │ unique_id_l │ unique_id_r │ first_name_l │ first_name_r │ gamma_first_name │    tf_first_name_l    │    tf_first_name_r    │   mw_first_name   │ mw_tf_adj_first_name │ surname_l │ surname_r │ gamma_surname │     tf_surname_l     │     tf_surname_r     │     mw_surname      │  mw_tf_adj_surname   │   dob_l    │   dob_r    │ gamma_dob │       mw_dob       │   city_l   │   city_r   │ gamma_city │      tf_city_l      │      tf_city_r      │      mw_city       │   mw_tf_adj_city    │         email_l          │         email_r          │ gamma_email │      tf_email_l       │      tf_email_r       │     mw_email      │ mw_tf_adj_email │ match_key │
│       double        │       double       │    int64    │    int64    │   varchar    │   varchar    │      int32       │        double         │        double         │      double       │        double        │  varchar  │  varchar  │     int32     │        double        │        double        │       double        │        double        │    date    │    date    │   int32   │       double       │  varchar   │  varchar   │   int32    │       double        │       double        │       double       │       double        │         varchar          │         varchar          │    int32    │        double         │        double         │      double       │     double      │  varchar  │
├─────────────────────┼────────────────────┼─────────────┼─────────────┼──────────────┼──────────────┼──────────────────┼───────────────────────┼───────────────────────┼───────────────────┼──────────────────────┼───────────┼───────────┼───────────────┼──────────────────────┼──────────────────────┼─────────────────────┼──────────────────────┼────────────┼────────────┼───────────┼────────────────────┼────────────┼────────────┼────────────┼─────────────────────┼─────────────────────┼────────────────────┼─────────────────────┼──────────────────────────┼──────────────────────────┼─────────────┼───────────────────────┼───────────────────────┼───────────────────┼─────────────────┼───────────┤
│  21.506834253252443 │ 0.9999996644187925 │          28 │          30 │ Thomas       │ Thomas       │                4 │  0.006016847172081829 │  0.006016847172081829 │ 6.309947619380791 │ 0.040393298600093885 │ Gabriel   │ Gabriel   │             4 │ 0.004884004884004884 │ 0.004884004884004884 │   6.580128192319569 │ -0.10616211501540729 │ 1976-09-15 │ 1976-09-15 │         2 │  7.715603236685051 │ London     │ London     │          1 │ 0.21279212792127922 │ 0.21279212792127922 │ 2.3410443182772083 │ -0.9401495213654414 │ gabriel.t54@nichols.info │ gabriel.t54@nlchois.info │           3 │ 0.0025348542458808617 │ 0.0012674271229404308 │ 7.952135813932513 │             0.0 │ 0         │
│ -2.7980062727528967 │  0.125710472856419 │          37 │         860 │ Theodore     │ Theodore     │                4 │  0.012033694344163659 │  0.012033694344163659 │ 6.309947619380791 │  -0.9596067013999061 │ Morris    │ Marshall  │             0 │ 0.004884004884004884 │ 0.007326007326007326 │ -2.1597988656731415 │                  0.0 │ 1978-08-19 │ 1972-07-25 │         0 │ -1.116036665159834 │ Birmingham │ Birmingham │          1 │ 0.04920049200492005 │ 0.04920049200492005 │ 2.3410443182772083 │  1.1725506113839206 │ t.m39@brooks-sawyer.com  │ NULL                     │          -1 │ 0.0063371356147021544 │                  NULL │               0.0 │             0.0 │ 0         │
│ -2.7980062727528967 │  0.125710472856419 │          39 │         860 │ Theodore     │ Theodore     │                4 │  0.012033694344163659 │  0.012033694344163659 │ 6.309947619380791 │  -0.9596067013999061 │ Morris    │ Marshall  │             0 │ 0.004884004884004884 │ 0.007326007326007326 │ -2.1597988656731415 │                  0.0 │ 1978-08-19 │ 1972-07-25 │         0 │ -1.116036665159834 │ Birmingham │ Birmingham │          1 │ 0.04920049200492005 │ 0.04920049200492005 │ 2.3410443182772083 │  1.1725506113839206 │ t.m39@brooks-sawyer.com  │ NULL                     │          -1 │ 0.0063371356147021544 │                  NULL │               0.0 │             0.0 │ 0         │
│ -2.7980062727528967 │  0.125710472856419 │          43 │         860 │ Theodore     │ Theodore     │                4 │  0.012033694344163659 │  0.012033694344163659 │ 6.309947619380791 │  -0.9596067013999061 │ Morris    │ Marshall  │             0 │ 0.004884004884004884 │ 0.007326007326007326 │ -2.1597988656731415 │                  0.0 │ 1978-08-19 │ 1972-07-25 │         0 │ -1.116036665159834 │ Birmingham │ Birmingham │          1 │ 0.04920049200492005 │ 0.04920049200492005 │ 2.3410443182772083 │  1.1725506113839206 │ t.m39@brooks-sawyer.com  │ NULL                     │          -1 │ 0.0063371356147021544 │                  NULL │               0.0 │             0.0 │ 0         │
│   6.523156368990757 │ 0.9892443210014298 │          47 │          49 │ Erin         │ Erin         │                4 │ 0.0048134777376654635 │ 0.0048134777376654635 │ 6.309947619380791 │   0.3623213934874556 │ Rogers    │ NULL      │            -1 │ 0.006105006105006105 │                 NULL │                 0.0 │                  0.0 │ 2010-01-02 │ 2010-03-01 │         0 │ -1.116036665159834 │ London     │ London     │          1 │ 0.21279212792127922 │ 0.21279212792127922 │ 2.3410443182772083 │ -0.9401495213654414 │ e.rogers3@hopkins.org    │ e.rogers3@honkips.org    │           3 │ 0.0025348542458808617 │ 0.0012674271229404308 │ 7.952135813932513 │             0.0 │ 0         │
└─────────────────────┴────────────────────┴─────────────┴─────────────┴──────────────┴──────────────┴──────────────────┴───────────────────────┴───────────────────────┴───────────────────┴──────────────────────┴───────────┴───────────┴───────────────┴──────────────────────┴──────────────────────┴─────────────────────┴──────────────────────┴────────────┴────────────┴───────────┴────────────────────┴────────────┴────────────┴────────────┴─────────────────────┴─────────────────────┴────────────────────┴─────────────────────┴──────────────────────────┴──────────────────────────┴─────────────┴───────────────────────┴───────────────────────┴───────────────────┴─────────────────┴───────────┘

Clustering

The result of linker.inference.predict() is a list of pairwise record comparisons and their associated scores. For instance, if we have input records A, B, C and D, it could be represented conceptually as:

A -> B with score 0.9
B -> C with score 0.95
C -> D with score 0.1
D -> E with score 0.99

Often, an alternative representation of this result is more useful, where each row is an input record, and where records link, they are assigned to the same cluster.

With a score threshold of 0.5, the above data could be represented conceptually as:

ID, Cluster ID
A,  1
B,  1
C,  1
D,  2
E,  2

The algorithm that converts between the pairwise results and the clusters is called connected components, and it is included in Splink. You can use it as follows:

clusters = linker.clustering.cluster_pairwise_predictions_at_threshold(
    df_predictions, threshold_match_probability=0.2
)
clusters.as_duckdbpyrelation().sort("cluster").limit(10).show(max_width=10000)
Completed iteration 1, num edges remaining to process: 380


Completed iteration 2, num edges remaining to process: 24


Completed iteration 3, num edges remaining to process: 0


┌────────────┬───────────┬────────────┬─────────┬────────────┬────────────┬───────────────────────────┬─────────┐
│ cluster_id │ unique_id │ first_name │ surname │    dob     │    city    │           email           │ cluster │
│   int64    │   int64   │  varchar   │ varchar │    date    │  varchar   │          varchar          │  int64  │
├────────────┼───────────┼────────────┼─────────┼────────────┼────────────┼───────────────────────────┼─────────┤
│          0 │         1 │ Robert     │ Allen   │ 1971-05-24 │ NULL       │ roberta25@smith.net       │       0 │
│          0 │         3 │ Robert     │ Alen    │ 1971-06-24 │ Lonon      │ NULL                      │       0 │
│          0 │         0 │ Robert     │ Alan    │ 1971-06-24 │ NULL       │ robert255@smith.net       │       0 │
│          0 │         2 │ Rob        │ Allen   │ 1971-06-24 │ London     │ roberta25@smith.net       │       0 │
│          4 │         4 │ Grace      │ NULL    │ 1997-04-26 │ Hull       │ grace.kelly52@jones.com   │       1 │
│          5 │         5 │ Grace      │ Kelly   │ 1991-04-26 │ NULL       │ grace.kelly52@jones.com   │       1 │
│          6 │         6 │ Logan      │ pMurphy │ 1973-08-01 │ NULL       │ NULL                      │       2 │
│          7 │         8 │ NULL       │ Dean    │ 2015-03-03 │ NULL       │ NULL                      │       3 │
│          7 │        10 │ NULL       │ Dean    │ 2015-03-03 │ Portsmouth │ evied56@harris-bailey.net │       3 │
│          7 │         9 │ Evie       │ Dean    │ 2015-03-03 │ Pootsmruth │ evihd56@earris-bailey.net │       3 │
└────────────┴───────────┴────────────┴─────────┴────────────┴────────────┴───────────────────────────┴─────────┘
  10 rows                                                                                             8 columns

The estimated cluster id is the cluster_id column, the true cluster id (which we only know because this data is labelled) is the cluster column.

Note that the value in the column isn't meaninful, all that matters is that it creates the right grouping. For instance, the three Evie Dean records are all assigned to estimated cluster_id 7, correctly grouping them together. It does not matter that the value used to group them together in the cluster column (the true label) is 3.

We can see the model has done a reasonable but imperfect job of estimating the true cluster. This is a simple model trained on a small and very messy dataset, so it is not surprising that accuracy is not better..

Further Reading

For more on the prediction tools in Splink, please refer to the Prediction API documentation.

Next steps

Now we have made predictions with a model, we can move on to visualising it to understand how it is working.