Skip to content

7. Evaluation

Evaluation of prediction results

Open In Colab

In the previous tutorial, we looked at various ways to visualise the results of our model. These are useful for evaluating a linkage pipeline because they allow us to understand how our model works and verify that it is doing something sensible. They can also be useful to identify examples where the model is not performing as expected.

In addition to these spot checks, Splink also has functions to perform more formal accuracy analysis. These functions allow you to understand the likely prevalence of false positives and false negatives in your linkage models.

They rely on the existence of a sample of labelled (ground truth) matches, which may have been produced (for example) by human beings. For the accuracy analysis to be unbiased, the sample should be representative of the overall dataset.

# Rerun our predictions to we're ready to view the charts
from splink import DuckDBAPI, Linker, splink_datasets

db_api = DuckDBAPI()
df = splink_datasets.fake_1000
df_sdf = db_api.register(df)
df_sdf.as_duckdbpyrelation().limit(5).show()
┌───────────┬────────────┬─────────┬────────────┬─────────┬─────────────────────────┬─────────┐
│ unique_id │ first_name │ surname │    dob     │  city   │          email          │ cluster │
│   int64   │  varchar   │ varchar │    date    │ varchar │         varchar         │  int64  │
├───────────┼────────────┼─────────┼────────────┼─────────┼─────────────────────────┼─────────┤
│         0 │ Robert     │ Alan    │ 1971-06-24 │ NULL    │ robert255@smith.net     │       0 │
│         1 │ Robert     │ Allen   │ 1971-05-24 │ NULL    │ roberta25@smith.net     │       0 │
│         2 │ Rob        │ Allen   │ 1971-06-24 │ London  │ roberta25@smith.net     │       0 │
│         3 │ Robert     │ Alen    │ 1971-06-24 │ Lonon   │ NULL                    │       0 │
│         4 │ Grace      │ NULL    │ 1997-04-26 │ Hull    │ grace.kelly52@jones.com │       1 │
└───────────┴────────────┴─────────┴────────────┴─────────┴─────────────────────────┴─────────┘
import urllib.request
from pathlib import Path


def get_settings_text() -> str:
    # assumes cwd is repo root
    local_path = Path.cwd() / "docs" / "demos" / "demo_settings" / "saved_model_from_demo.json"

    if local_path.exists():
        return local_path.read_text()

    # fallback location for settings - the file as it is on master, for e.g. colab use
    # TODO: update ref
    url = "https://raw.githubusercontent.com/moj-analytical-services/splink/master/docs/demos/demo_settings/saved_model_from_demo.json"
    with urllib.request.urlopen(url) as u:
        return u.read().decode()
import json

from splink import block_on


settings = json.loads(get_settings_text())

# The data quality is very poor in this dataset, so we need looser blocking rules
# to achieve decent recall
settings["blocking_rules_to_generate_predictions"] = [
    block_on("first_name"),
    block_on("city"),
    block_on("email"),
    block_on("dob"),
]

linker = Linker(df_sdf, settings)
df_predictions = linker.inference.predict(threshold_match_probability=0.01)
Blocking time: 0.01 seconds


Predict time (post-blocking): 0.10 seconds

Load in labels

The labels file contains a list of pairwise comparisons which represent matches and non-matches.

The required format of the labels file is described here.

from splink.datasets import splink_dataset_labels
from splink.internals.misc import show

df_labels = splink_dataset_labels.fake_1000_labels
df_labels_sdf = db_api.register(df_labels)
labels_table = linker.table_management.register_labels_table(df_labels_sdf)
show(df_labels, rows=5)
┌─────────────┬──────────────────┬─────────────┬──────────────────┬──────────────────────┐
│ unique_id_l │ source_dataset_l │ unique_id_r │ source_dataset_r │ clerical_match_score │
│    int64    │     varchar      │    int64    │     varchar      │        double        │
├─────────────┼──────────────────┼─────────────┼──────────────────┼──────────────────────┤
│           0 │ fake_1000        │           1 │ fake_1000        │                  1.0 │
│           0 │ fake_1000        │           2 │ fake_1000        │                  1.0 │
│           0 │ fake_1000        │           3 │ fake_1000        │                  1.0 │
│           0 │ fake_1000        │           4 │ fake_1000        │                  0.0 │
│           0 │ fake_1000        │           5 │ fake_1000        │                  0.0 │
└─────────────┴──────────────────┴─────────────┴──────────────────┴──────────────────────┘

View examples of false positives and false negatives

splink_df = linker.evaluation.prediction_errors_from_labels_table(
    labels_table, include_false_negatives=True, include_false_positives=False
)
false_negatives = splink_df.as_record_list(limit=5)
linker.visualisations.waterfall_chart(false_negatives)

False positives

# Note I've picked a threshold match probability of 0.01 here because otherwise
# in this simple example there are no false positives
splink_df = linker.evaluation.prediction_errors_from_labels_table(
    labels_table, include_false_negatives=False, include_false_positives=True, threshold_match_probability=0.01
)
false_postives = splink_df.as_record_list(limit=5)
linker.visualisations.waterfall_chart(false_postives)

Threshold Selection chart

Splink includes an interactive dashboard that shows key accuracy statistics:

linker.evaluation.accuracy_analysis_from_labels_table(
    labels_table, output_type="threshold_selection", add_metrics=["f1"]
)

Precision recall chart

A Precision-Recall chart shows how the number of false positives and false negatives varies depending on the match threshold chosen. The match threshold is the match weight chosen as a cutoff for which pairwise comparisons to accept as matches.

linker.evaluation.accuracy_analysis_from_labels_table(labels_table, output_type="precision_recall")

Truth table

Finally, Splink can also report the underlying table used to construct the precision-recall curves.

roc_table = linker.evaluation.accuracy_analysis_from_labels_table(
    labels_table, output_type="table"
)
roc_table.as_duckdbpyrelation().limit(5).show(max_width=10000)
┌─────────────────────┬────────────────────────┬───────────────────────┬────────┬────────┬────────┬────────┬────────┬────────┬────────────────────┬────────────┬────────────────────┬────────────────────┬──────────────────────┬─────────────────────┬────────────────────┬────────────────────┬────────────────────┬────────────────────┬────────────────────┬────────────────────┬────────────────────┬────────────────────┬────────────────────┬────────────────────┐
│   truth_threshold   │   match_probability    │ total_clerical_labels │   p    │   n    │   tp   │   tn   │   fp   │   fn   │       P_rate       │   N_rate   │      tp_rate       │      tn_rate       │       fp_rate        │       fn_rate       │     precision      │       recall       │    specificity     │        npv         │      accuracy      │         f1         │         f2         │        f0_5        │         p4         │        phi         │
│       double        │         double         │        double         │ double │ double │ double │ double │ double │ double │       double       │   float    │       double       │       double       │        double        │       double        │       double       │       double       │       double       │       double       │       double       │       double       │       double       │       double       │       double       │       double       │
├─────────────────────┼────────────────────────┼───────────────────────┼────────┼────────┼────────┼────────┼────────┼────────┼────────────────────┼────────────┼────────────────────┼────────────────────┼──────────────────────┼─────────────────────┼────────────────────┼────────────────────┼────────────────────┼────────────────────┼────────────────────┼────────────────────┼────────────────────┼────────────────────┼────────────────────┼────────────────────┤
│  -18.80000028014183 │ 2.1909630111470102e-06 │                3176.0 │ 2031.0 │ 1145.0 │ 1709.0 │ 1103.0 │   42.0 │  322.0 │ 0.6394836272040302 │ 0.36051637 │ 0.8414574101427869 │ 0.9633187772925764 │  0.03668122270742358 │  0.1585425898572132 │ 0.9760137064534552 │ 0.8414574101427869 │ 0.9633187772925764 │ 0.7740350877192983 │ 0.8853904282115869 │ 0.9037546271813856 │ 0.8653164556962025 │ 0.9457664637520753 │ 0.8804756275225732 │ 0.7769307620147627 │
│ -16.700000248849392 │  9.392796608724036e-06 │                3176.0 │ 2031.0 │ 1145.0 │ 1709.0 │ 1119.0 │   26.0 │  322.0 │ 0.6394836272040302 │ 0.36051637 │ 0.8414574101427869 │  0.977292576419214 │ 0.022707423580786028 │  0.1585425898572132 │  0.985014409221902 │ 0.8414574101427869 │  0.977292576419214 │ 0.7765440666204025 │ 0.8904282115869018 │ 0.9075942644715879 │ 0.8667207627548433 │ 0.9525136551109129 │ 0.8860103770975539 │ 0.7896366201374305 │
│ -12.700000189244747 │ 0.00015026358101882152 │                3176.0 │ 2031.0 │ 1145.0 │ 1709.0 │ 1125.0 │   20.0 │  322.0 │ 0.6394836272040302 │ 0.36051637 │ 0.8414574101427869 │  0.982532751091703 │ 0.017467248908296942 │  0.1585425898572132 │ 0.9884326200115674 │ 0.8414574101427869 │  0.982532751091703 │ 0.7774706288873532 │ 0.8923173803526449 │ 0.9090425531914894 │ 0.8672485537399777 │  0.955068738124511 │ 0.8880763922377238 │ 0.7944159751353451 │
│ -12.500000186264515 │ 0.00017260367204143044 │                3176.0 │ 2031.0 │ 1145.0 │ 1708.0 │ 1125.0 │   20.0 │  323.0 │ 0.6394836272040302 │ 0.36051637 │ 0.8409650418513048 │  0.982532751091703 │ 0.017467248908296942 │ 0.15903495814869523 │ 0.9884259259259259 │ 0.8409650418513048 │  0.982532751091703 │ 0.7769337016574586 │ 0.8920025188916877 │ 0.9087523277467412 │ 0.8668290702395453 │ 0.9549368220954937 │  0.887762700545028 │ 0.7938966961277768 │
│ -12.300000183284283 │ 0.00019826446591752426 │                3176.0 │ 2031.0 │ 1145.0 │ 1705.0 │ 1132.0 │   13.0 │  326.0 │ 0.6394836272040302 │ 0.36051637 │ 0.8394879369768586 │  0.988646288209607 │ 0.011353711790393014 │  0.1605120630231413 │ 0.9924330616996507 │ 0.8394879369768586 │  0.988646288209607 │ 0.7764060356652949 │ 0.8932619647355163 │ 0.9095758869031741 │ 0.8661857346067873 │ 0.9575424014377176 │ 0.8892254223487883 │ 0.7979360689863448 │
└─────────────────────┴────────────────────────┴───────────────────────┴────────┴────────┴────────┴────────┴────────┴────────┴────────────────────┴────────────┴────────────────────┴────────────────────┴──────────────────────┴─────────────────────┴────────────────────┴────────────────────┴────────────────────┴────────────────────┴────────────────────┴────────────────────┴────────────────────┴────────────────────┴────────────────────┴────────────────────┘

Unlinkables chart

Finally, it can be interesting to analyse whether your dataset contains any 'unlinkable' records.

'Unlinkable records' are records with such poor data quality they don't even link to themselves at a high enough probability to be accepted as matches

For example, in a typical linkage problem, a 'John Smith' record with nulls for their address and postcode may be unlinkable. By 'unlinkable' we don't mean there are no matches; rather, we mean it is not possible to determine whether there are matches.UnicodeTranslateError

A high proportion of unlinkable records is an indication of poor quality in the input dataset

linker.evaluation.unlinkables_chart()

For this dataset and this trained model, we can see that most records are (theoretically) linkable: At a match weight 6, around around 99% of records could be linked to themselves.

Further Reading

For more on the quality assurance tools in Splink, please refer to the Evaluation API documentation.

📊 For more on the charts used in this tutorial, please refer to the Charts Gallery.

For more on the Evaluation Metrics used in this tutorial, please refer to the Edge Metrics guide.