Skip to content

Evaluation from ground truth column

Evaluation when you have fully labelled data

In this example, our data contains a fully-populated ground-truth column called cluster that enables us to perform accuracy analysis of the final model

Open In Colab

from splink import splink_datasets
from splink.internals.misc import show

df = splink_datasets.fake_1000
show(df, rows=5)
┌───────────┬────────────┬─────────┬────────────┬─────────┬─────────────────────────┬─────────┐
│ unique_id │ first_name │ surname │    dob     │  city   │          email          │ cluster │
│   int64   │  varchar   │ varchar │    date    │ varchar │         varchar         │  int64  │
├───────────┼────────────┼─────────┼────────────┼─────────┼─────────────────────────┼─────────┤
│         0 │ Robert     │ Alan    │ 1971-06-24 │ NULL    │ robert255@smith.net     │       0 │
│         1 │ Robert     │ Allen   │ 1971-05-24 │ NULL    │ roberta25@smith.net     │       0 │
│         2 │ Rob        │ Allen   │ 1971-06-24 │ London  │ roberta25@smith.net     │       0 │
│         3 │ Robert     │ Alen    │ 1971-06-24 │ Lonon   │ NULL                    │       0 │
│         4 │ Grace      │ NULL    │ 1997-04-26 │ Hull    │ grace.kelly52@jones.com │       1 │
└───────────┴────────────┴─────────┴────────────┴─────────┴─────────────────────────┴─────────┘
from splink import SettingsCreator, Linker, block_on, DuckDBAPI

import splink.comparison_library as cl

settings = SettingsCreator(
    link_type="dedupe_only",
    blocking_rules_to_generate_predictions=[
        block_on("first_name"),
        block_on("surname"),
        block_on("dob"),
        block_on("email"),
    ],
    comparisons=[
        cl.ForenameSurnameComparison("first_name", "surname"),
        cl.DateOfBirthComparison(
            "dob",
            input_is_string=False,
        ),
        cl.ExactMatch("city").configure(term_frequency_adjustments=True),
        cl.EmailComparison("email"),
    ],
    retain_intermediate_calculation_columns=True,
)
db_api = DuckDBAPI()
df_sdf = db_api.register(df)
linker = Linker(df_sdf, settings)
deterministic_rules = [
    "l.first_name = r.first_name and levenshtein(r.dob::VARCHAR, l.dob::VARCHAR) <= 1",
    "l.surname = r.surname and levenshtein(r.dob::VARCHAR, l.dob::VARCHAR) <= 1",
    "l.first_name = r.first_name and levenshtein(r.surname, l.surname) <= 2",
    "l.email = r.email",
]

linker.training.estimate_probability_two_random_records_match(
    deterministic_rules, recall=0.7
)
Probability two random records match is estimated to be  0.00333.
This means that amongst all possible pairwise record comparisons, one in 300.13 are expected to match.  With 499,500 total possible comparisons, we expect a total of around 1,664.29 matching pairs
linker.training.estimate_u_using_random_sampling(max_pairs=1e6, seed=5)
You are using the default value for `max_pairs`, which may be too small and thus lead to inaccurate estimates for your model's u-parameters. Consider increasing to 1e8 or 1e9, which will result in more accurate estimates, but with a longer run time.


----- Estimating u probabilities using random sampling -----


Estimating u with: max_pairs = 1,000,000, min_count_per_level = 100, num_chunks = 10



Estimating u for: first_name_surname (Comparison 1 of 4)


  Running probe chunk (~1.00% of max_pairs)


  Min u_count: 2 for comparison level (Jaro-Winkler distance of first_name >= 0.88) AND (Jaro-Winkler distance of surname >= 0.88) (cvv=3)


  Probe did not converge; restarting with normal chunking



  Running chunk 1/10


  Count of 6 for level (Jaro-Winkler distance of first_name >= 0.88) AND (Jaro-Winkler distance of surname >= 0.88) (cvv=3). Chunk took 0.0 seconds.


  Min u_count not hit, continuing.


  Running chunk 2/10


  Count of 7 for level (Jaro-Winkler distance of first_name >= 0.88) AND (Jaro-Winkler distance of surname >= 0.88) (cvv=3). Chunk took 0.0 seconds.


  Min u_count not hit, continuing.


  Running chunk 3/10


  Count of 15 for level (Jaro-Winkler distance of first_name >= 0.88) AND (Jaro-Winkler distance of surname >= 0.88) (cvv=3). Chunk took 0.0 seconds.


  Min u_count not hit, continuing.


  Running chunk 4/10


  Count of 18 for level (Jaro-Winkler distance of first_name >= 0.88) AND (Jaro-Winkler distance of surname >= 0.88) (cvv=3). Chunk took 0.0 seconds.


  Min u_count not hit, continuing.


  Running chunk 5/10


  Count of 29 for level (Jaro-Winkler distance of first_name >= 0.88) AND (Jaro-Winkler distance of surname >= 0.88) (cvv=3). Chunk took 0.0 seconds.


  Min u_count not hit, continuing.


  Running chunk 6/10


  Count of 39 for level (Jaro-Winkler distance of first_name >= 0.88) AND (Jaro-Winkler distance of surname >= 0.88) (cvv=3). Chunk took 0.0 seconds.


  Min u_count not hit, continuing.


  Running chunk 7/10


  Count of 51 for level (Jaro-Winkler distance of first_name >= 0.88) AND (Jaro-Winkler distance of surname >= 0.88) (cvv=3). Chunk took 0.0 seconds.


  Min u_count not hit, continuing.


  Running chunk 8/10


  Count of 66 for level (Jaro-Winkler distance of first_name >= 0.88) AND (Jaro-Winkler distance of surname >= 0.88) (cvv=3). Chunk took 0.0 seconds.


  Min u_count not hit, continuing.


  Running chunk 9/10


  Count of 70 for level (Jaro-Winkler distance of first_name >= 0.88) AND (Jaro-Winkler distance of surname >= 0.88) (cvv=3). Chunk took 0.0 seconds.


  Min u_count not hit, continuing.


  Running chunk 10/10


  Count of 72 for level (Jaro-Winkler distance of first_name >= 0.88) AND (Jaro-Winkler distance of surname >= 0.88) (cvv=3). Chunk took 0.0 seconds.


  Min u_count not hit, continuing.



Estimating u for: dob (Comparison 2 of 4)


  Running probe chunk (~1.00% of max_pairs)


  Min u_count: 6 for comparison level Abs date difference <= 1 month (cvv=3)


  Probe did not converge; restarting with normal chunking



  Running chunk 1/10


  Count of 92 for level Levenshtein distance <= 1 (cvv=4). Chunk took 0.1 seconds.


  Min u_count not hit, continuing.


  Running chunk 2/10


  Count of 154 for level Levenshtein distance <= 1 (cvv=4). Chunk took 0.1 seconds.


  Exiting early since min count of 154 exceeds min_count_per_level = 100



Estimating u for: city (Comparison 3 of 4)


  Running probe chunk (~1.00% of max_pairs)


  Min u_count: 414 for comparison level Exact match on city (cvv=1)


  Exiting early since min count of 414 exceeds min_count_per_level = 100



Estimating u for: email (Comparison 4 of 4)


  Running probe chunk (~1.00% of max_pairs)


  Min u_count: 1 for comparison level Jaro-Winkler >0.88 on username (cvv=1)


  Probe did not converge; restarting with normal chunking



  Running chunk 1/10


  Count of 13 for level Jaro-Winkler >0.88 on username (cvv=1). Chunk took 0.1 seconds.


  Min u_count not hit, continuing.


  Running chunk 2/10


  Count of 22 for level Jaro-Winkler >0.88 on username (cvv=1). Chunk took 0.1 seconds.


  Min u_count not hit, continuing.


  Running chunk 3/10


  Count of 35 for level Jaro-Winkler >0.88 on username (cvv=1). Chunk took 0.1 seconds.


  Min u_count not hit, continuing.


  Running chunk 4/10


  Count of 61 for level Jaro-Winkler >0.88 on username (cvv=1). Chunk took 0.1 seconds.


  Min u_count not hit, continuing.


  Running chunk 5/10


  Count of 99 for level Jaro-Winkler >0.88 on username (cvv=1). Chunk took 0.1 seconds.


  Min u_count not hit, continuing.


  Running chunk 6/10


  Count of 113 for level Jaro-Winkler >0.88 on username (cvv=1). Chunk took 0.1 seconds.


  Exiting early since min count of 113 exceeds min_count_per_level = 100



Estimated u probabilities using random sampling



Your model is not yet fully trained. Missing estimates for:
    - first_name_surname (no m values are trained).
    - dob (no m values are trained).
    - city (no m values are trained).
    - email (no m values are trained).
session_dob = linker.training.estimate_parameters_using_expectation_maximisation(
    block_on("dob"), estimate_without_term_frequencies=True
)
session_email = linker.training.estimate_parameters_using_expectation_maximisation(
    block_on("email"), estimate_without_term_frequencies=True
)
session_dob = linker.training.estimate_parameters_using_expectation_maximisation(
    block_on("first_name", "surname"), estimate_without_term_frequencies=True
)
----- Starting EM training session -----



[EM sampling] max_pairs is None — no sampling will be applied


Estimating the m probabilities of the model by blocking on:
l."dob" = r."dob"

Parameter estimates will be made for the following comparison(s):
    - first_name_surname
    - city
    - email

Parameter estimates cannot be made for the following comparison(s) since they are used in the blocking rules: 
    - dob





WARNING:
Level Jaro-Winkler >0.88 on username on comparison email not observed in dataset, unable to train m value



Iteration 1: Largest change in params was -0.747 in the m_probability of first_name_surname, level `(Exact match on first_name) AND (Exact match on surname)`


Iteration 2: Largest change in params was 0.196 in probability_two_random_records_match


Iteration 3: Largest change in params was 0.0589 in probability_two_random_records_match


Iteration 4: Largest change in params was 0.0212 in probability_two_random_records_match


Iteration 5: Largest change in params was 0.00826 in probability_two_random_records_match


Iteration 6: Largest change in params was 0.0033 in probability_two_random_records_match


Iteration 7: Largest change in params was 0.00133 in probability_two_random_records_match


Iteration 8: Largest change in params was 0.00054 in probability_two_random_records_match


Iteration 9: Largest change in params was 0.000219 in probability_two_random_records_match


Iteration 10: Largest change in params was 8.88e-05 in probability_two_random_records_match



EM converged after 10 iterations


m probability not trained for email - Jaro-Winkler >0.88 on username (comparison vector value: 1). This usually means the comparison level was never observed in the training data.



Your model is not yet fully trained. Missing estimates for:
    - dob (no m values are trained).
    - email (some m values are not trained).



----- Starting EM training session -----



[EM sampling] max_pairs is None — no sampling will be applied


Estimating the m probabilities of the model by blocking on:
l."email" = r."email"

Parameter estimates will be made for the following comparison(s):
    - first_name_surname
    - dob
    - city

Parameter estimates cannot be made for the following comparison(s) since they are used in the blocking rules: 
    - email





Iteration 1: Largest change in params was -0.435 in the m_probability of dob, level `Exact match on date of birth`


Iteration 2: Largest change in params was 0.121 in probability_two_random_records_match


Iteration 3: Largest change in params was 0.0311 in probability_two_random_records_match


Iteration 4: Largest change in params was 0.0113 in probability_two_random_records_match


Iteration 5: Largest change in params was 0.00517 in probability_two_random_records_match


Iteration 6: Largest change in params was 0.00277 in probability_two_random_records_match


Iteration 7: Largest change in params was 0.00165 in probability_two_random_records_match


Iteration 8: Largest change in params was 0.00106 in probability_two_random_records_match


Iteration 9: Largest change in params was 0.000712 in probability_two_random_records_match


Iteration 10: Largest change in params was 0.000496 in probability_two_random_records_match


Iteration 11: Largest change in params was 0.000354 in probability_two_random_records_match


Iteration 12: Largest change in params was 0.000257 in probability_two_random_records_match


Iteration 13: Largest change in params was 0.000189 in probability_two_random_records_match


Iteration 14: Largest change in params was 0.000141 in probability_two_random_records_match


Iteration 15: Largest change in params was 0.000105 in probability_two_random_records_match


Iteration 16: Largest change in params was 7.94e-05 in probability_two_random_records_match



EM converged after 16 iterations



Your model is not yet fully trained. Missing estimates for:
    - email (some m values are not trained).



----- Starting EM training session -----



[EM sampling] max_pairs is None — no sampling will be applied


Estimating the m probabilities of the model by blocking on:
(l."first_name" = r."first_name") AND (l."surname" = r."surname")

Parameter estimates will be made for the following comparison(s):
    - dob
    - city
    - email

Parameter estimates cannot be made for the following comparison(s) since they are used in the blocking rules: 
    - first_name_surname





WARNING:
Level Jaro-Winkler >0.88 on username on comparison email not observed in dataset, unable to train m value



Iteration 1: Largest change in params was 0.465 in probability_two_random_records_match


Iteration 2: Largest change in params was 0.0496 in probability_two_random_records_match


Iteration 3: Largest change in params was 0.0095 in probability_two_random_records_match


Iteration 4: Largest change in params was 0.0019 in probability_two_random_records_match


Iteration 5: Largest change in params was 0.000385 in probability_two_random_records_match


Iteration 6: Largest change in params was 8.55e-05 in probability_two_random_records_match



EM converged after 6 iterations


m probability not trained for email - Jaro-Winkler >0.88 on username (comparison vector value: 1). This usually means the comparison level was never observed in the training data.



Your model is not yet fully trained. Missing estimates for:
    - email (some m values are not trained).
linker.evaluation.accuracy_analysis_from_labels_column(
    "cluster", output_type="table"
).as_duckdbpyrelation().limit(5).show(max_width=10000)
Blocking time: 0.01 seconds


Predict time (post-blocking): 0.10 seconds



 -- WARNING --
You have called predict(), but there are some parameter estimates which have neither been estimated or specified in your settings dictionary.  To produce predictions the following untrained parameters will use default values.
Comparison: 'email':
    m values not fully trained


┌─────────────────────┬────────────────────────┬───────────────────────┬────────┬──────────┬────────┬──────────┬────────┬────────┬──────────────────────┬────────────────────┬────────────────────┬────────────────────┬───────────────────────┬─────────────────────┬─────────────────────┬────────────────────┬────────────────────┬────────────────────┬────────────────────┬────────────────────┬────────────────────┬─────────────────────┬────────────────────┬────────────────────┐
│   truth_threshold   │   match_probability    │ total_clerical_labels │   p    │    n     │   tp   │    tn    │   fp   │   fn   │        P_rate        │       N_rate       │      tp_rate       │      tn_rate       │        fp_rate        │       fn_rate       │      precision      │       recall       │    specificity     │        npv         │      accuracy      │         f1         │         f2         │        f0_5         │         p4         │        phi         │
│       double        │         double         │        double         │ double │  double  │ double │  double  │ double │ double │        double        │       double       │       double       │       double       │        double         │       double        │       double        │       double       │       double       │       double       │       double       │       double       │       double       │       double        │       double       │       double       │
├─────────────────────┼────────────────────────┼───────────────────────┼────────┼──────────┼────────┼──────────┼────────┼────────┼──────────────────────┼────────────────────┼────────────────────┼────────────────────┼───────────────────────┼─────────────────────┼─────────────────────┼────────────────────┼────────────────────┼────────────────────┼────────────────────┼────────────────────┼────────────────────┼─────────────────────┼────────────────────┼────────────────────┤
│ -28.500000424683094 │  2.634177249574674e-09 │              499500.0 │ 2031.0 │ 497469.0 │ 1650.0 │ 495130.0 │ 2339.0 │  381.0 │ 0.004066066066066066 │ 0.9959339339339339 │ 0.8124076809453471 │ 0.9952981994857971 │ 0.0047018005142028954 │ 0.18759231905465287 │ 0.41363750313361747 │ 0.8124076809453471 │ 0.9952981994857971 │ 0.9992310967869533 │ 0.9945545545545545 │ 0.5481727574750831 │  0.681086436060431 │ 0.45866459109356755 │ 0.7074664508208482 │ 0.5774741518035404 │
│ -28.400000423192978 │ 2.8232412740932453e-09 │              499500.0 │ 2031.0 │ 497469.0 │ 1650.0 │ 495225.0 │ 2244.0 │  381.0 │ 0.004066066066066066 │ 0.9959339339339339 │ 0.8124076809453471 │ 0.9954891661590973 │  0.004510833840902649 │ 0.18759231905465287 │   0.423728813559322 │ 0.8124076809453471 │ 0.9954891661590973 │ 0.9992312441737994 │ 0.9947447447447447 │ 0.5569620253164557 │ 0.6864702945581628 │  0.4685636394615778 │ 0.7147694968503172 │ 0.5845580356933797 │
│  -27.80000041425228 │   4.27923359068787e-09 │              499500.0 │ 2031.0 │ 497469.0 │ 1650.0 │ 495311.0 │ 2158.0 │  381.0 │ 0.004066066066066066 │ 0.9959339339339339 │ 0.8124076809453471 │ 0.9956620412528218 │  0.004337958747178216 │ 0.18759231905465287 │  0.4332983193277311 │ 0.8124076809453471 │ 0.9956620412528218 │ 0.9992313775489619 │  0.994916916916917 │ 0.5651652680253468 │ 0.6914180355346966 │ 0.47790071250651683 │ 0.7215119201499676 │ 0.5911972192065922 │
│ -27.700000412762165 │  4.586369005821638e-09 │              499500.0 │ 2031.0 │ 497469.0 │ 1650.0 │ 495354.0 │ 2115.0 │  381.0 │ 0.004066066066066066 │ 0.9959339339339339 │ 0.8124076809453471 │  0.995748478799684 │  0.004251521200315999 │ 0.18759231905465287 │ 0.43824701195219123 │ 0.8124076809453471 │  0.995748478799684 │ 0.9992314442191896 │  0.995003003003003 │ 0.5693581780538303 │ 0.6939187484229119 │  0.4827101983500088 │ 0.7249310558395574 │ 0.5946014708278546 │
│  -27.60000041127205 │  4.915548593297618e-09 │              499500.0 │ 2031.0 │ 497469.0 │ 1650.0 │ 495386.0 │ 2083.0 │  381.0 │ 0.004066066066066066 │ 0.9959339339339339 │ 0.8124076809453471 │ 0.9958128044159535 │ 0.0041871955840464435 │ 0.18759231905465287 │ 0.44200375033485134 │ 0.8124076809453471 │ 0.9958128044159535 │ 0.9992314938267371 │  0.995067067067067 │ 0.5725190839694656 │ 0.6957915155604284 │  0.4863526498850439 │ 0.7274966332137337 │ 0.5971728084857775 │
└─────────────────────┴────────────────────────┴───────────────────────┴────────┴──────────┴────────┴──────────┴────────┴────────┴──────────────────────┴────────────────────┴────────────────────┴────────────────────┴───────────────────────┴─────────────────────┴─────────────────────┴────────────────────┴────────────────────┴────────────────────┴────────────────────┴────────────────────┴────────────────────┴─────────────────────┴────────────────────┴────────────────────┘
linker.evaluation.accuracy_analysis_from_labels_column("cluster", output_type="precision_recall")
Blocking time: 0.01 seconds


Predict time (post-blocking): 0.10 seconds



 -- WARNING --
You have called predict(), but there are some parameter estimates which have neither been estimated or specified in your settings dictionary.  To produce predictions the following untrained parameters will use default values.
Comparison: 'email':
    m values not fully trained
linker.evaluation.accuracy_analysis_from_labels_column(
    "cluster",
    output_type="threshold_selection",
    threshold_match_probability=0.5,
    add_metrics=["f1"],
)
Blocking time: 0.01 seconds


Predict time (post-blocking): 0.09 seconds



 -- WARNING --
You have called predict(), but there are some parameter estimates which have neither been estimated or specified in your settings dictionary.  To produce predictions the following untrained parameters will use default values.
Comparison: 'email':
    m values not fully trained
# Plot some false positives
linker.evaluation.prediction_errors_from_labels_column(
    "cluster", include_false_negatives=True, include_false_positives=True
).as_duckdbpyrelation().limit(5).show(max_width=10000)
Blocking time: 0.01 seconds


Predict time (post-blocking): 0.12 seconds



 -- WARNING --
You have called predict(), but there are some parameter estimates which have neither been estimated or specified in your settings dictionary.  To produce predictions the following untrained parameters will use default values.
Comparison: 'email':
    m values not fully trained


┌──────────────────────┬─────────────────────────┬─────────────────────┬──────────────────────┬─────────────┬─────────────┬───────────┬───────────┬──────────────┬──────────────┬──────────────────────────┬──────────────────────┬──────────────────────┬───────────────────────┬───────────────────────┬───────────────────────┬──────────────────────────────┬────────────┬────────────┬───────────┬────────────────────┬────────────┬─────────┬────────────┬──────────────────────┬──────────────────────┬─────────────────────┬─────────────────────┬───────────────┬───────────────────────────────┬─────────────┬───────────────────────┬───────────────────────┬──────────┬─────────────────┬───────────┬───────────┬───────────┐
│ clerical_match_score │ found_by_blocking_rules │    match_weight     │  match_probability   │ unique_id_l │ unique_id_r │ surname_l │ surname_r │ first_name_l │ first_name_r │ gamma_first_name_surname │     tf_surname_l     │     tf_surname_r     │    tf_first_name_l    │    tf_first_name_r    │ mw_first_name_surname │ mw_tf_adj_first_name_surname │   dob_l    │   dob_r    │ gamma_dob │       mw_dob       │   city_l   │ city_r  │ gamma_city │      tf_city_l       │      tf_city_r       │       mw_city       │   mw_tf_adj_city    │    email_l    │            email_r            │ gamma_email │      tf_email_l       │      tf_email_r       │ mw_email │ mw_tf_adj_email │ cluster_l │ cluster_r │ match_key │
│        double        │         boolean         │       double        │        double        │    int64    │    int64    │  varchar  │  varchar  │   varchar    │   varchar    │          int32           │        double        │        double        │        double         │        double         │        double         │            double            │    date    │    date    │   int32   │       double       │  varchar   │ varchar │   int32    │        double        │        double        │       double        │       double        │    varchar    │            varchar            │    int32    │        double         │        double         │  double  │     double      │   int64   │   int64   │  varchar  │
├──────────────────────┼─────────────────────────┼─────────────────────┼──────────────────────┼─────────────┼─────────────┼───────────┼───────────┼──────────────┼──────────────┼──────────────────────────┼──────────────────────┼──────────────────────┼───────────────────────┼───────────────────────┼───────────────────────┼──────────────────────────────┼────────────┼────────────┼───────────┼────────────────────┼────────────┼─────────┼────────────┼──────────────────────┼──────────────────────┼─────────────────────┼─────────────────────┼───────────────┼───────────────────────────────┼─────────────┼───────────────────────┼───────────────────────┼──────────┼─────────────────┼───────────┼───────────┼───────────┤
│                  1.0 │ true                    │ -2.6683489893903696 │  0.13592473325083895 │         177 │         178 │ Ellis     │ Eislls    │ Ellie        │ Ellie        │                        1 │ 0.003663003663003663 │ 0.001221001221001221 │ 0.0024067388688327317 │ 0.0024067388688327317 │      5.42416367418533 │           0.5613535502605806 │ 2015-11-06 │ 2005-12-08 │         1 │ -0.429243420096612 │ London     │ NULL    │         -1 │  0.21279212792127922 │                 NULL │                 0.0 │                 0.0 │ NULL          │ elliee@oguzmanoc.m            │          -1 │                  NULL │ 0.0012674271229404308 │      0.0 │             0.0 │        48 │        48 │ 0         │
│                  1.0 │ true                    │ -1.8255353171145083 │  0.22005177871263146 │         239 │         242 │ Shah      │ Shah      │ Freya        │ Freya        │                        6 │ 0.003663003663003663 │ 0.003663003663003663 │  0.006016847172081829 │  0.006016847172081829 │     7.865170525138507 │                          0.0 │ 1972-01-17 │ 1970-12-17 │         1 │ -0.429243420096612 │ London     │ Lonnod  │          0 │  0.21279212792127922 │ 0.007380073800738007 │ -1.0368396284167354 │                 0.0 │ f.s@flynn.com │ NULL                          │          -1 │  0.005069708491761723 │                  NULL │      0.0 │             0.0 │        62 │        62 │ 0         │
│                  1.0 │ true                    │  -3.990277084277732 │  0.05919775688883316 │         248 │         249 │ Williams  │ NULL      │ Joshua       │ Joshua       │                        1 │ 0.009768009768009768 │                 NULL │  0.006016847172081829 │  0.006016847172081829 │      5.42416367418533 │           -0.760574544626782 │ 2012-08-05 │ 2003-09-07 │         1 │ -0.429243420096612 │ Portsmouth │ NULL    │         -1 │ 0.017220172201722016 │                 NULL │                 0.0 │                 0.0 │ NULL          │ j.williams@levine-johnson.com │          -1 │                  NULL │ 0.0076045627376425855 │      0.0 │             0.0 │        64 │        64 │ 0         │
│                  1.0 │ true                    │  -7.585014473304475 │ 0.005181161464484199 │         376 │         380 │ NULL      │ Taylor    │ Eliza        │ Eliza        │                        1 │                 NULL │ 0.017094017094017096 │  0.006016847172081829 │  0.006016847172081829 │      5.42416367418533 │           -0.760574544626782 │ 1983-01-12 │ 1993-02-08 │         0 │ -4.023980809123355 │ Glasgow    │ NULL    │         -1 │ 0.008610086100861008 │                 NULL │                 0.0 │                 0.0 │ NULL          │ elizataylor@marshall.com      │          -1 │                  NULL │ 0.0063371356147021544 │      0.0 │             0.0 │        98 │        98 │ 0         │
│                  1.0 │ true                    │  -2.970396153671388 │  0.11315399061466086 │         434 │         435 │ Kaur      │ Kaur      │ Hugo         │ Hugo         │                        6 │ 0.009768009768009768 │ 0.009768009768009768 │ 0.0048134777376654635 │ 0.0048134777376654635 │     7.865170525138507 │                          0.0 │ 1983-04-20 │ 1973-03-23 │         0 │ -4.023980809123355 │ London     │ London  │          1 │  0.21279212792127922 │  0.21279212792127922 │   2.353186445418569 │ -0.9401495213654414 │ h.k@sharp.com │ NULL                          │          -1 │ 0.0038022813688212928 │                  NULL │      0.0 │             0.0 │       112 │       112 │ 0         │
└──────────────────────┴─────────────────────────┴─────────────────────┴──────────────────────┴─────────────┴─────────────┴───────────┴───────────┴──────────────┴──────────────┴──────────────────────────┴──────────────────────┴──────────────────────┴───────────────────────┴───────────────────────┴───────────────────────┴──────────────────────────────┴────────────┴────────────┴───────────┴────────────────────┴────────────┴─────────┴────────────┴──────────────────────┴──────────────────────┴─────────────────────┴─────────────────────┴───────────────┴───────────────────────────────┴─────────────┴───────────────────────┴───────────────────────┴──────────┴─────────────────┴───────────┴───────────┴───────────┘
records = linker.evaluation.prediction_errors_from_labels_column(
    "cluster", include_false_negatives=True, include_false_positives=True
).as_record_list(limit=5)

linker.visualisations.waterfall_chart(records)
Blocking time: 0.02 seconds


Predict time (post-blocking): 0.17 seconds



 -- WARNING --
You have called predict(), but there are some parameter estimates which have neither been estimated or specified in your settings dictionary.  To produce predictions the following untrained parameters will use default values.
Comparison: 'email':
    m values not fully trained