Skip to content

Deduplicate 50k rows historical persons

Linking a dataset of real historical persons¶

In this example, we deduplicate a more realistic dataset. The data is based on historical persons scraped from wikidata. Duplicate records are introduced with a variety of errors introduced.

Note, as explained in the backends topic guide, SQLite does not natively support string fuzzy matching functions such as damareau-levenshtein and jaro-winkler (as used in this example). Instead, these have been imported as python User Defined Functions (UDFs). One drawback of python UDFs is that they are considerably slower than native-SQL comparisons. As such, if you are hitting issues with large run times, consider switching to DuckDB (or some other backend).

Open In Colab

# Uncomment and run this cell if you're running in Google Colab.
# !pip install "splink[altair,igraph,pyarrow]"
# !pip install rapidfuzz
from splink import splink_datasets

# reduce size of dataset to make things run faster
df = splink_datasets.historical_50k.slice(0, 5000)
from splink.backends.sqlite import SQLiteAPI
from splink.exploratory import profile_columns

db_api = SQLiteAPI()
df_sdf = db_api.register(df)
profile_columns(
    df_sdf, column_expressions=["first_name", "postcode_fake", "substr(dob, 1,4)"]
)
from splink import block_on
from splink.blocking_analysis import (
    chart_comparisons_from_blocking_rules,
)

blocking_rules =  [block_on("first_name", "surname"),
        block_on("surname", "dob"),
        block_on("first_name", "dob"),
        block_on("postcode_fake", "first_name")]



chart_comparisons_from_blocking_rules(
    df_sdf,
    blocking_rules=blocking_rules,
    link_type="dedupe_only",
    record_sample_proportion=0.2,
)
/home/runner/work/splink/splink/splink/internals/blocking_analysis.py:668: UserWarning: The sampled blocking analysis estimate for blocking rule '(l."first_name" = r."first_name") AND (l."surname" = r."surname")' is based on 456 sampled pairwise comparisons. This is below the recommended minimum of 1,000, so the estimate may be unstable. Increase record_sample_proportion for a more stable estimate.
  return _cumulative_comparisons_to_be_scored_from_blocking_rules(
/home/runner/work/splink/splink/splink/internals/blocking_analysis.py:668: UserWarning: The sampled blocking analysis estimate for blocking rule '(l."surname" = r."surname") AND (l."dob" = r."dob")' is based on 91 sampled pairwise comparisons. This is below the recommended minimum of 1,000, so the estimate may be unstable. Increase record_sample_proportion for a more stable estimate.
  return _cumulative_comparisons_to_be_scored_from_blocking_rules(
/home/runner/work/splink/splink/splink/internals/blocking_analysis.py:668: UserWarning: The sampled blocking analysis estimate for blocking rule '(l."first_name" = r."first_name") AND (l."dob" = r."dob")' is based on 37 sampled pairwise comparisons. This is below the recommended minimum of 1,000, so the estimate may be unstable. Increase record_sample_proportion for a more stable estimate.
  return _cumulative_comparisons_to_be_scored_from_blocking_rules(
/home/runner/work/splink/splink/splink/internals/blocking_analysis.py:668: UserWarning: The sampled blocking analysis estimate for blocking rule '(l."postcode_fake" = r."postcode_fake") AND (l."first_name" = r."first_name")' is based on 26 sampled pairwise comparisons. This is below the recommended minimum of 1,000, so the estimate may be unstable. Increase record_sample_proportion for a more stable estimate.
  return _cumulative_comparisons_to_be_scored_from_blocking_rules(
import splink.comparison_library as cl
from splink import Linker

settings = {
    "link_type": "dedupe_only",
    "blocking_rules_to_generate_predictions": [
        block_on("first_name", "surname"),
        block_on("surname", "dob"),
        block_on("first_name", "dob"),
        block_on("postcode_fake", "first_name"),

    ],
    "comparisons": [
        cl.NameComparison("first_name"),
        cl.NameComparison("surname"),
        cl.DamerauLevenshteinAtThresholds("dob", [1, 2]).configure(
            term_frequency_adjustments=True
        ),
        cl.DamerauLevenshteinAtThresholds("postcode_fake", [1, 2]),
        cl.ExactMatch("birth_place").configure(term_frequency_adjustments=True),
        cl.ExactMatch(
            "occupation",
        ).configure(term_frequency_adjustments=True),
    ],
    "retain_matching_columns": True,
    "retain_intermediate_calculation_columns": True,
    "max_iterations": 10,
    "em_convergence": 0.01,
}

linker = Linker(df_sdf, settings)
linker.training.estimate_probability_two_random_records_match(
    [
        block_on("first_name", "surname", "dob"),
        block_on("substr(first_name,1,2)", "surname", "substr(postcode_fake,1,2)"),
        block_on("dob", "postcode_fake"),
    ],
    recall=0.6,
)
Probability two random records match is estimated to be  0.00139.
This means that amongst all possible pairwise record comparisons, one in 721.98 are expected to match.  With 12,497,500 total possible comparisons, we expect a total of around 17,310.00 matching pairs
linker.training.estimate_u_using_random_sampling(max_pairs=1e6)
You are using the default value for `max_pairs`, which may be too small and thus lead to inaccurate estimates for your model's u-parameters. Consider increasing to 1e8 or 1e9, which will result in more accurate estimates, but with a longer run time.


----- Estimating u probabilities using random sampling -----


Estimating u with: max_pairs = 1,000,000, min_count_per_level = 100, num_chunks = 10



Estimating u for: first_name (Comparison 1 of 6)


  Running probe chunk (~1.00% of max_pairs)


  Min u_count: 0 for comparison level Jaro-Winkler distance of first_name >= 0.7 (cvv=1)


  Probe did not converge; restarting with normal chunking



  Running chunk 1/10


  Count of 0 for level Jaro-Winkler distance of first_name >= 0.7 (cvv=1). Chunk took 0.4 seconds.


  Min u_count not hit, continuing.


  Running chunk 2/10


  Count of 0 for level Jaro-Winkler distance of first_name >= 0.7 (cvv=1). Chunk took 0.4 seconds.


  Min u_count not hit, continuing.


  Running chunk 3/10


  Count of 0 for level Jaro-Winkler distance of first_name >= 0.7 (cvv=1). Chunk took 0.4 seconds.


  Min u_count not hit, continuing.


  Running chunk 4/10


  Count of 0 for level Jaro-Winkler distance of first_name >= 0.7 (cvv=1). Chunk took 0.4 seconds.


  Min u_count not hit, continuing.


  Running chunk 5/10


  Count of 0 for level Jaro-Winkler distance of first_name >= 0.7 (cvv=1). Chunk took 0.4 seconds.


  Min u_count not hit, continuing.


  Running chunk 6/10


  Count of 0 for level Jaro-Winkler distance of first_name >= 0.7 (cvv=1). Chunk took 0.4 seconds.


  Min u_count not hit, continuing.


  Running chunk 7/10


  Count of 0 for level Jaro-Winkler distance of first_name >= 0.7 (cvv=1). Chunk took 0.3 seconds.


  Min u_count not hit, continuing.


  Running chunk 8/10


  Count of 0 for level Jaro-Winkler distance of first_name >= 0.7 (cvv=1). Chunk took 0.3 seconds.


  Min u_count not hit, continuing.


  Running chunk 9/10


  Count of 0 for level Jaro-Winkler distance of first_name >= 0.7 (cvv=1). Chunk took 0.3 seconds.


  Min u_count not hit, continuing.


  Running chunk 10/10


  Count of 0 for level Jaro-Winkler distance of first_name >= 0.7 (cvv=1). Chunk took 0.3 seconds.


  Min u_count not hit, continuing.



Estimating u for: surname (Comparison 2 of 6)


  Running probe chunk (~1.00% of max_pairs)


  Min u_count: 0 for comparison level Jaro-Winkler distance of surname >= 0.88 (cvv=2)


  Probe did not converge; restarting with normal chunking



  Running chunk 1/10


  Count of 0 for level Jaro-Winkler distance of surname >= 0.88 (cvv=2). Chunk took 0.4 seconds.


  Min u_count not hit, continuing.


  Running chunk 2/10


  Count of 0 for level Jaro-Winkler distance of surname >= 0.88 (cvv=2). Chunk took 0.4 seconds.


  Min u_count not hit, continuing.


  Running chunk 3/10


  Count of 0 for level Jaro-Winkler distance of surname >= 0.88 (cvv=2). Chunk took 0.3 seconds.


  Min u_count not hit, continuing.


  Running chunk 4/10


  Count of 0 for level Jaro-Winkler distance of surname >= 0.88 (cvv=2). Chunk took 0.4 seconds.


  Min u_count not hit, continuing.


  Running chunk 5/10


  Count of 0 for level Jaro-Winkler distance of surname >= 0.88 (cvv=2). Chunk took 0.4 seconds.


  Min u_count not hit, continuing.


  Running chunk 6/10


  Count of 0 for level Jaro-Winkler distance of surname >= 0.88 (cvv=2). Chunk took 0.5 seconds.


  Min u_count not hit, continuing.


  Running chunk 7/10


  Count of 0 for level Jaro-Winkler distance of surname >= 0.88 (cvv=2). Chunk took 0.4 seconds.


  Min u_count not hit, continuing.


  Running chunk 8/10


  Count of 0 for level Jaro-Winkler distance of surname >= 0.88 (cvv=2). Chunk took 0.5 seconds.


  Min u_count not hit, continuing.


  Running chunk 9/10


  Count of 0 for level Jaro-Winkler distance of surname >= 0.88 (cvv=2). Chunk took 0.4 seconds.


  Min u_count not hit, continuing.


  Running chunk 10/10


  Count of 0 for level Jaro-Winkler distance of surname >= 0.88 (cvv=2). Chunk took 0.3 seconds.


  Min u_count not hit, continuing.



Estimating u for: dob (Comparison 3 of 6)


  Running probe chunk (~1.00% of max_pairs)


  Min u_count: 2 for comparison level Exact match on dob (cvv=3)


  Probe did not converge; restarting with normal chunking



  Running chunk 1/10


  Count of 255 for level Exact match on dob (cvv=3). Chunk took 0.2 seconds.


  Exiting early since min count of 255 exceeds min_count_per_level = 100



Estimating u for: postcode_fake (Comparison 4 of 6)


  Running probe chunk (~1.00% of max_pairs)


  Min u_count: 0 for comparison level Damerau-Levenshtein distance of postcode_fake <= 2 (cvv=1)


  Probe did not converge; restarting with normal chunking



  Running chunk 1/10


  Count of 19 for level Damerau-Levenshtein distance of postcode_fake <= 1 (cvv=2). Chunk took 0.3 seconds.


  Min u_count not hit, continuing.


  Running chunk 2/10


  Count of 47 for level Damerau-Levenshtein distance of postcode_fake <= 1 (cvv=2). Chunk took 0.2 seconds.


  Min u_count not hit, continuing.


  Running chunk 3/10


  Count of 71 for level Damerau-Levenshtein distance of postcode_fake <= 1 (cvv=2). Chunk took 0.2 seconds.


  Min u_count not hit, continuing.


  Running chunk 4/10


  Count of 98 for level Damerau-Levenshtein distance of postcode_fake <= 1 (cvv=2). Chunk took 0.3 seconds.


  Min u_count not hit, continuing.


  Running chunk 5/10


  Count of 118 for level Damerau-Levenshtein distance of postcode_fake <= 1 (cvv=2). Chunk took 0.3 seconds.


  Exiting early since min count of 118 exceeds min_count_per_level = 100



Estimating u for: birth_place (Comparison 5 of 6)


  Running probe chunk (~1.00% of max_pairs)


  Min u_count: 23 for comparison level Exact match on birth_place (cvv=1)


  Probe did not converge; restarting with normal chunking



  Running chunk 1/10


  Count of 516 for level Exact match on birth_place (cvv=1). Chunk took 0.2 seconds.


  Exiting early since min count of 516 exceeds min_count_per_level = 100



Estimating u for: occupation (Comparison 6 of 6)


  Running probe chunk (~1.00% of max_pairs)


  Min u_count: 42 for comparison level Exact match on occupation (cvv=1)


  Probe did not converge; restarting with normal chunking



  Running chunk 1/10


  Count of 1,334 for level Exact match on occupation (cvv=1). Chunk took 0.2 seconds.


  Exiting early since min count of 1,334 exceeds min_count_per_level = 100



Estimated u probabilities using random sampling



Your model is not yet fully trained. Missing estimates for:
    - first_name (some u values are not trained, no m values are trained).
    - surname (some u values are not trained, no m values are trained).
    - dob (no m values are trained).
    - postcode_fake (no m values are trained).
    - birth_place (no m values are trained).
    - occupation (no m values are trained).
training_blocking_rule = "l.first_name = r.first_name and l.surname = r.surname"
training_session_names = linker.training.estimate_parameters_using_expectation_maximisation(
    training_blocking_rule, estimate_without_term_frequencies=True
)
----- Starting EM training session -----



[EM sampling] max_pairs is None — no sampling will be applied


Estimating the m probabilities of the model by blocking on:
l.first_name = r.first_name and l.surname = r.surname

Parameter estimates will be made for the following comparison(s):
    - dob
    - postcode_fake
    - birth_place
    - occupation

Parameter estimates cannot be made for the following comparison(s) since they are used in the blocking rules: 
    - first_name
    - surname





Iteration 1: Largest change in params was -0.332 in the m_probability of dob, level `Exact match on dob`


Iteration 2: Largest change in params was 0.0146 in the m_probability of postcode_fake, level `All other comparisons`


Iteration 3: Largest change in params was -0.00126 in the m_probability of dob, level `All other comparisons`



EM converged after 3 iterations



Your model is not yet fully trained. Missing estimates for:
    - first_name (some u values are not trained, no m values are trained).
    - surname (some u values are not trained, no m values are trained).
training_blocking_rule = "l.dob = r.dob"
training_session_dob = linker.training.estimate_parameters_using_expectation_maximisation(
    training_blocking_rule, estimate_without_term_frequencies=True
)
----- Starting EM training session -----



[EM sampling] max_pairs is None — no sampling will be applied


Estimating the m probabilities of the model by blocking on:
l.dob = r.dob

Parameter estimates will be made for the following comparison(s):
    - first_name
    - surname
    - postcode_fake
    - birth_place
    - occupation

Parameter estimates cannot be made for the following comparison(s) since they are used in the blocking rules: 
    - dob





WARNING:
Level Jaro-Winkler distance of first_name >= 0.88 on comparison first_name not observed in dataset, unable to train m value



WARNING:
Level Jaro-Winkler distance of first_name >= 0.7 on comparison first_name not observed in dataset, unable to train m value



WARNING:
Level Jaro-Winkler distance of surname >= 0.88 on comparison surname not observed in dataset, unable to train m value



WARNING:
Level Jaro-Winkler distance of surname >= 0.7 on comparison surname not observed in dataset, unable to train m value



Iteration 1: Largest change in params was -0.327 in the m_probability of first_name, level `Exact match on first_name`


Iteration 2: Largest change in params was -0.0634 in the m_probability of first_name, level `Exact match on first_name`


Iteration 3: Largest change in params was -0.0113 in the m_probability of surname, level `Exact match on surname`


Iteration 4: Largest change in params was -0.00245 in the m_probability of surname, level `Exact match on surname`



EM converged after 4 iterations


m probability not trained for first_name - Jaro-Winkler distance of first_name >= 0.88 (comparison vector value: 2). This usually means the comparison level was never observed in the training data.


m probability not trained for first_name - Jaro-Winkler distance of first_name >= 0.7 (comparison vector value: 1). This usually means the comparison level was never observed in the training data.


m probability not trained for surname - Jaro-Winkler distance of surname >= 0.88 (comparison vector value: 2). This usually means the comparison level was never observed in the training data.


m probability not trained for surname - Jaro-Winkler distance of surname >= 0.7 (comparison vector value: 1). This usually means the comparison level was never observed in the training data.



Your model is not yet fully trained. Missing estimates for:
    - first_name (some u values are not trained, some m values are not trained).
    - surname (some u values are not trained, some m values are not trained).

The final match weights can be viewed in the match weights chart:

linker.visualisations.match_weights_chart()
linker.evaluation.unlinkables_chart()
df_predict = linker.inference.predict()
df_predict.as_record_list(limit=5)
Blocking time: 0.04 seconds


Predict time (post-blocking): 1.03 seconds



 -- WARNING --
You have called predict(), but there are some parameter estimates which have neither been estimated or specified in your settings dictionary.  To produce predictions the following untrained parameters will use default values.
Comparison: 'first_name':
    m values not fully trained
Comparison: 'first_name':
    u values not fully trained
Comparison: 'surname':
    m values not fully trained
Comparison: 'surname':
    u values not fully trained





[{'match_weight': 21.280574169596314,
  'match_probability': 0.9999996074376185,
  'unique_id_l': 'Q2296770-1',
  'unique_id_r': 'Q2296770-10',
  'first_name_l': 'thomas',
  'first_name_r': 'thomas',
  'gamma_first_name': 4,
  'tf_first_name_l': 0.03162530024019215,
  'tf_first_name_r': 0.03162530024019215,
  'mw_first_name': 5.30326619708439,
  'mw_tf_adj_first_name': -1.1819306724792442,
  'surname_l': 'chudleigh',
  'surname_r': 'chudleigh',
  'gamma_surname': 4,
  'tf_surname_l': 0.003051438535309503,
  'tf_surname_r': 0.003051438535309503,
  'mw_surname': 8.485191337054985,
  'mw_tf_adj_surname': -0.4906377373049988,
  'dob_l': '1630-08-01',
  'dob_r': None,
  'gamma_dob': -1,
  'tf_dob_l': 0.0026028110359187923,
  'tf_dob_r': None,
  'mw_dob': 0.0,
  'mw_tf_adj_dob': 0.0,
  'postcode_fake_l': 'tq13 8df',
  'postcode_fake_r': 'tq13 8jr',
  'gamma_postcode_fake': 1,
  'mw_postcode_fake': 6.5463847481237085,
  'birth_place_l': 'devon',
  'birth_place_r': 'devon',
  'gamma_birth_place': 1,
  'tf_birth_place_l': 0.0020901068276823038,
  'tf_birth_place_r': 0.0020901068276823038,
  'mw_birth_place': 7.091563178607089,
  'mw_tf_adj_birth_place': 1.5850404197957184,
  'occupation_l': 'politician',
  'occupation_r': 'politician',
  'gamma_occupation': 1,
  'tf_occupation_l': 0.08285243198680957,
  'tf_occupation_r': 0.08285243198680957,
  'mw_occupation': 4.304679225460623,
  'mw_tf_adj_occupation': -0.86916406864826,
  'match_key': '0'},
 {'match_weight': 17.857878514129162,
  'match_probability': 0.9999957903975828,
  'unique_id_l': 'Q2296770-1',
  'unique_id_r': 'Q2296770-14',
  'first_name_l': 'thomas',
  'first_name_r': 'thomas',
  'gamma_first_name': 4,
  'tf_first_name_l': 0.03162530024019215,
  'tf_first_name_r': 0.03162530024019215,
  'mw_first_name': 5.30326619708439,
  'mw_tf_adj_first_name': -1.1819306724792442,
  'surname_l': 'chudleigh',
  'surname_r': 'chudleigh',
  'gamma_surname': 4,
  'tf_surname_l': 0.003051438535309503,
  'tf_surname_r': 0.003051438535309503,
  'mw_surname': 8.485191337054985,
  'mw_tf_adj_surname': -0.4906377373049988,
  'dob_l': '1630-08-01',
  'dob_r': '1638-08-01',
  'gamma_dob': 2,
  'tf_dob_l': 0.0026028110359187923,
  'tf_dob_r': 0.0002602811035918792,
  'mw_dob': 3.8748091455805356,
  'mw_tf_adj_dob': 0.0,
  'postcode_fake_l': 'tq13 8df',
  'postcode_fake_r': 'tq1w 8df',
  'gamma_postcode_fake': 2,
  'mw_postcode_fake': 7.925483545478829,
  'birth_place_l': 'devon',
  'birth_place_r': None,
  'gamma_birth_place': -1,
  'tf_birth_place_l': 0.0020901068276823038,
  'tf_birth_place_r': None,
  'mw_birth_place': 0.0,
  'mw_tf_adj_birth_place': 0.0,
  'occupation_l': 'politician',
  'occupation_r': 'politician',
  'gamma_occupation': 1,
  'tf_occupation_l': 0.08285243198680957,
  'tf_occupation_r': 0.08285243198680957,
  'mw_occupation': 4.304679225460623,
  'mw_tf_adj_occupation': -0.86916406864826,
  'match_key': '0'},
 {'match_weight': 31.72687176849857,
  'match_probability': 0.9999999997186415,
  'unique_id_l': 'Q2296770-1',
  'unique_id_r': 'Q2296770-2',
  'first_name_l': 'thomas',
  'first_name_r': 'thomas',
  'gamma_first_name': 4,
  'tf_first_name_l': 0.03162530024019215,
  'tf_first_name_r': 0.03162530024019215,
  'mw_first_name': 5.30326619708439,
  'mw_tf_adj_first_name': -1.1819306724792442,
  'surname_l': 'chudleigh',
  'surname_r': 'chudleigh',
  'gamma_surname': 4,
  'tf_surname_l': 0.003051438535309503,
  'tf_surname_r': 0.003051438535309503,
  'mw_surname': 8.485191337054985,
  'mw_tf_adj_surname': -0.4906377373049988,
  'dob_l': '1630-08-01',
  'dob_r': '1630-08-01',
  'gamma_dob': 3,
  'tf_dob_l': 0.0026028110359187923,
  'tf_dob_r': 0.0026028110359187923,
  'mw_dob': 7.24479377100702,
  'mw_tf_adj_dob': 0.6310155993490687,
  'postcode_fake_l': 'tq13 8df',
  'postcode_fake_r': 'tq13 8df',
  'gamma_postcode_fake': 3,
  'mw_postcode_fake': 9.11687297666987,
  'birth_place_l': 'devon',
  'birth_place_r': 'devon',
  'gamma_birth_place': 1,
  'tf_birth_place_l': 0.0020901068276823038,
  'tf_birth_place_r': 0.0020901068276823038,
  'mw_birth_place': 7.091563178607089,
  'mw_tf_adj_birth_place': 1.5850404197957184,
  'occupation_l': 'politician',
  'occupation_r': 'politician',
  'gamma_occupation': 1,
  'tf_occupation_l': 0.08285243198680957,
  'tf_occupation_r': 0.08285243198680957,
  'mw_occupation': 4.304679225460623,
  'mw_tf_adj_occupation': -0.86916406864826,
  'match_key': '0'},
 {'match_weight': 29.156383539952408,
  'match_probability': 0.9999999983287016,
  'unique_id_l': 'Q2296770-1',
  'unique_id_r': 'Q2296770-4',
  'first_name_l': 'thomas',
  'first_name_r': 'thomas',
  'gamma_first_name': 4,
  'tf_first_name_l': 0.03162530024019215,
  'tf_first_name_r': 0.03162530024019215,
  'mw_first_name': 5.30326619708439,
  'mw_tf_adj_first_name': -1.1819306724792442,
  'surname_l': 'chudleigh',
  'surname_r': 'chudleigh',
  'gamma_surname': 4,
  'tf_surname_l': 0.003051438535309503,
  'tf_surname_r': 0.003051438535309503,
  'mw_surname': 8.485191337054985,
  'mw_tf_adj_surname': -0.4906377373049988,
  'dob_l': '1630-08-01',
  'dob_r': '1630-08-01',
  'gamma_dob': 3,
  'tf_dob_l': 0.0026028110359187923,
  'tf_dob_r': 0.0026028110359187923,
  'mw_dob': 7.24479377100702,
  'mw_tf_adj_dob': 0.6310155993490687,
  'postcode_fake_l': 'tq13 8df',
  'postcode_fake_r': 'tq13 8hu',
  'gamma_postcode_fake': 1,
  'mw_postcode_fake': 6.5463847481237085,
  'birth_place_l': 'devon',
  'birth_place_r': 'devon',
  'gamma_birth_place': 1,
  'tf_birth_place_l': 0.0020901068276823038,
  'tf_birth_place_r': 0.0020901068276823038,
  'mw_birth_place': 7.091563178607089,
  'mw_tf_adj_birth_place': 1.5850404197957184,
  'occupation_l': 'politician',
  'occupation_r': 'politician',
  'gamma_occupation': 1,
  'tf_occupation_l': 0.08285243198680957,
  'tf_occupation_r': 0.08285243198680957,
  'mw_occupation': 4.304679225460623,
  'mw_tf_adj_occupation': -0.86916406864826,
  'match_key': '0'},
 {'match_weight': 31.72687176849857,
  'match_probability': 0.9999999997186415,
  'unique_id_l': 'Q2296770-1',
  'unique_id_r': 'Q2296770-5',
  'first_name_l': 'thomas',
  'first_name_r': 'thomas',
  'gamma_first_name': 4,
  'tf_first_name_l': 0.03162530024019215,
  'tf_first_name_r': 0.03162530024019215,
  'mw_first_name': 5.30326619708439,
  'mw_tf_adj_first_name': -1.1819306724792442,
  'surname_l': 'chudleigh',
  'surname_r': 'chudleigh',
  'gamma_surname': 4,
  'tf_surname_l': 0.003051438535309503,
  'tf_surname_r': 0.003051438535309503,
  'mw_surname': 8.485191337054985,
  'mw_tf_adj_surname': -0.4906377373049988,
  'dob_l': '1630-08-01',
  'dob_r': '1630-08-01',
  'gamma_dob': 3,
  'tf_dob_l': 0.0026028110359187923,
  'tf_dob_r': 0.0026028110359187923,
  'mw_dob': 7.24479377100702,
  'mw_tf_adj_dob': 0.6310155993490687,
  'postcode_fake_l': 'tq13 8df',
  'postcode_fake_r': 'tq13 8df',
  'gamma_postcode_fake': 3,
  'mw_postcode_fake': 9.11687297666987,
  'birth_place_l': 'devon',
  'birth_place_r': 'devon',
  'gamma_birth_place': 1,
  'tf_birth_place_l': 0.0020901068276823038,
  'tf_birth_place_r': 0.0020901068276823038,
  'mw_birth_place': 7.091563178607089,
  'mw_tf_adj_birth_place': 1.5850404197957184,
  'occupation_l': 'politician',
  'occupation_r': 'politician',
  'gamma_occupation': 1,
  'tf_occupation_l': 0.08285243198680957,
  'tf_occupation_r': 0.08285243198680957,
  'mw_occupation': 4.304679225460623,
  'mw_tf_adj_occupation': -0.86916406864826,
  'match_key': '0'}]

You can also view rows in this dataset as a waterfall chart as follows:

records_to_plot = df_predict.as_record_list(limit=5)
linker.visualisations.waterfall_chart(records_to_plot, filter_nulls=False)
clusters = linker.clustering.cluster_pairwise_predictions_at_threshold(
    df_predict, threshold_match_probability=0.95
)
Completed iteration 1, num edges remaining to process: 1302


Completed iteration 2, num edges remaining to process: 94


Completed iteration 3, num edges remaining to process: 10


Completed iteration 4, num edges remaining to process: 0
linker.visualisations.cluster_studio_dashboard(
    df_predict,
    clusters,
    "dashboards/50k_cluster.html",
    sampling_method="by_cluster_size",
    overwrite=True,
)

from IPython.display import IFrame

IFrame(src="./dashboards/50k_cluster.html", width="100%", height=1200)

linker.evaluation.accuracy_analysis_from_labels_column(
    "cluster", output_type="roc", match_weight_round_to_nearest=0.02
)
Blocking time: 0.05 seconds


Predict time (post-blocking): 3.80 seconds



 -- WARNING --
You have called predict(), but there are some parameter estimates which have neither been estimated or specified in your settings dictionary.  To produce predictions the following untrained parameters will use default values.
Comparison: 'first_name':
    m values not fully trained
Comparison: 'first_name':
    u values not fully trained
Comparison: 'surname':
    m values not fully trained
Comparison: 'surname':
    u values not fully trained
records = linker.evaluation.prediction_errors_from_labels_column(
    "cluster",
    threshold_match_probability=0.999,
    include_false_negatives=False,
    include_false_positives=True,
).as_record_list()
linker.visualisations.waterfall_chart(records)
Blocking time: 0.09 seconds


Predict time (post-blocking): 5.79 seconds



 -- WARNING --
You have called predict(), but there are some parameter estimates which have neither been estimated or specified in your settings dictionary.  To produce predictions the following untrained parameters will use default values.
Comparison: 'first_name':
    m values not fully trained
Comparison: 'first_name':
    u values not fully trained
Comparison: 'surname':
    m values not fully trained
Comparison: 'surname':
    u values not fully trained
# Some of the false negatives will be because they weren't detected by the blocking rules
records = linker.evaluation.prediction_errors_from_labels_column(
    "cluster",
    threshold_match_probability=0.5,
    include_false_negatives=True,
    include_false_positives=False,
).as_record_list(limit=50)

linker.visualisations.waterfall_chart(records)
Blocking time: 0.10 seconds


Predict time (post-blocking): 6.00 seconds



 -- WARNING --
You have called predict(), but there are some parameter estimates which have neither been estimated or specified in your settings dictionary.  To produce predictions the following untrained parameters will use default values.
Comparison: 'first_name':
    m values not fully trained
Comparison: 'first_name':
    u values not fully trained
Comparison: 'surname':
    m values not fully trained
Comparison: 'surname':
    u values not fully trained