Linking two tables of persons
Linking without deduplication¶
A simple record linkage model using the link_only link type.
With link_only, only between-dataset record comparisons are generated. No within-dataset record comparisons are created, meaning that the model does not attempt to find within-dataset duplicates.
from splink import splink_datasets
from splink.internals.misc import show
import duckdb
df = splink_datasets.fake_1000
# Split a simple dataset into two, separate datasets which can be linked together.
rel = duckdb.sql("select * from df")
df_l = rel.filter("hash(unique_id) % 2 = 0").arrow().read_all()
df_r = rel.filter("hash(unique_id) % 2 = 1").arrow().read_all()
show(df_r, rows=2)
show(df_l, rows=2)
┌───────────┬────────────┬─────────┬────────────┬─────────┬─────────┬─────────┐
│ unique_id │ first_name │ surname │ dob │ city │ email │ cluster │
│ int64 │ varchar │ varchar │ date │ varchar │ varchar │ int64 │
├───────────┼────────────┼─────────┼────────────┼─────────┼─────────┼─────────┤
│ 6 │ Logan │ pMurphy │ 1973-08-01 │ NULL │ NULL │ 2 │
│ 8 │ NULL │ Dean │ 2015-03-03 │ NULL │ NULL │ 3 │
└───────────┴────────────┴─────────┴────────────┴─────────┴─────────┴─────────┘
┌───────────┬────────────┬─────────┬────────────┬─────────┬─────────────────────┬─────────┐
│ unique_id │ first_name │ surname │ dob │ city │ email │ cluster │
│ int64 │ varchar │ varchar │ date │ varchar │ varchar │ int64 │
├───────────┼────────────┼─────────┼────────────┼─────────┼─────────────────────┼─────────┤
│ 0 │ Robert │ Alan │ 1971-06-24 │ NULL │ robert255@smith.net │ 0 │
│ 1 │ Robert │ Allen │ 1971-05-24 │ NULL │ roberta25@smith.net │ 0 │
└───────────┴────────────┴─────────┴────────────┴─────────┴─────────────────────┴─────────┘
import splink.comparison_library as cl
from splink import DuckDBAPI, Linker, SettingsCreator, block_on
settings = SettingsCreator(
link_type="link_only",
blocking_rules_to_generate_predictions=[
block_on("first_name"),
block_on("surname"),
],
comparisons=[
cl.NameComparison(
"first_name",
),
cl.NameComparison("surname"),
cl.DateOfBirthComparison(
"dob",
input_is_string=False,
),
cl.ExactMatch("city").configure(term_frequency_adjustments=True),
cl.EmailComparison("email"),
],
)
db_api = DuckDBAPI()
df_l_sdf = db_api.register(df_l, dataset_display_name="df_left")
df_r_sdf = db_api.register(df_r, dataset_display_name="df_right")
linker = Linker([df_l_sdf, df_r_sdf], settings)
from splink.exploratory import completeness_chart
db_api = DuckDBAPI()
df_l_sdf = db_api.register(df_l, dataset_display_name="df_left")
df_r_sdf = db_api.register(df_r, dataset_display_name="df_right")
completeness_chart(
[df_l_sdf, df_r_sdf],
cols=["first_name", "surname", "dob", "city", "email"],
table_names_for_chart=["df_left", "df_right"],
)
deterministic_rules = [
"l.first_name = r.first_name and levenshtein(r.dob::VARCHAR, l.dob::VARCHAR) <= 1",
"l.surname = r.surname and levenshtein(r.dob::VARCHAR, l.dob::VARCHAR) <= 1",
"l.first_name = r.first_name and levenshtein(r.surname, l.surname) <= 2",
block_on("email"),
]
linker.training.estimate_probability_two_random_records_match(deterministic_rules, recall=0.7)
Probability two random records match is estimated to be 0.0034.
This means that amongst all possible pairwise record comparisons, one in 293.93 are expected to match. With 249,424 total possible comparisons, we expect a total of around 848.57 matching pairs
linker.training.estimate_u_using_random_sampling(max_pairs=1e6, seed=1)
You are using the default value for `max_pairs`, which may be too small and thus lead to inaccurate estimates for your model's u-parameters. Consider increasing to 1e8 or 1e9, which will result in more accurate estimates, but with a longer run time.
----- Estimating u probabilities using random sampling -----
Estimating u with: max_pairs = 1,000,000, min_count_per_level = 100, num_chunks = 10
Estimating u for: first_name (Comparison 1 of 5)
Running probe chunk (~1.00% of max_pairs)
Min u_count: 9 for comparison level Jaro-Winkler distance of first_name >= 0.88 (cvv=2)
Probe did not converge; restarting with normal chunking
Running chunk 1/10
Count of 37 for level Jaro-Winkler distance of first_name >= 0.88 (cvv=2). Chunk took 0.0 seconds.
Min u_count not hit, continuing.
Running chunk 2/10
Count of 58 for level Jaro-Winkler distance of first_name >= 0.88 (cvv=2). Chunk took 0.0 seconds.
Min u_count not hit, continuing.
Running chunk 3/10
Count of 107 for level Jaro-Winkler distance of first_name >= 0.88 (cvv=2). Chunk took 0.0 seconds.
Exiting early since min count of 107 exceeds min_count_per_level = 100
Estimating u for: surname (Comparison 2 of 5)
Running probe chunk (~1.00% of max_pairs)
Min u_count: 2 for comparison level Jaro-Winkler distance of surname >= 0.88 (cvv=2)
Probe did not converge; restarting with normal chunking
Running chunk 1/10
Count of 27 for level Jaro-Winkler distance of surname >= 0.88 (cvv=2). Chunk took 0.0 seconds.
Min u_count not hit, continuing.
Running chunk 2/10
Count of 53 for level Jaro-Winkler distance of surname >= 0.88 (cvv=2). Chunk took 0.0 seconds.
Min u_count not hit, continuing.
Running chunk 3/10
Count of 63 for level Jaro-Winkler distance of surname >= 0.88 (cvv=2). Chunk took 0.0 seconds.
Min u_count not hit, continuing.
Running chunk 4/10
Count of 69 for level Jaro-Winkler distance of surname >= 0.88 (cvv=2). Chunk took 0.0 seconds.
Min u_count not hit, continuing.
Running chunk 5/10
Count of 91 for level Jaro-Winkler distance of surname >= 0.88 (cvv=2). Chunk took 0.0 seconds.
Min u_count not hit, continuing.
Running chunk 6/10
Count of 120 for level Jaro-Winkler distance of surname >= 0.88 (cvv=2). Chunk took 0.0 seconds.
Exiting early since min count of 120 exceeds min_count_per_level = 100
Estimating u for: dob (Comparison 3 of 5)
Running probe chunk (~1.00% of max_pairs)
Min u_count: 2 for comparison level Abs date difference <= 1 month (cvv=3)
Probe did not converge; restarting with normal chunking
Running chunk 1/10
Count of 34 for level Levenshtein distance <= 1 (cvv=4). Chunk took 0.0 seconds.
Min u_count not hit, continuing.
Running chunk 2/10
Count of 58 for level Levenshtein distance <= 1 (cvv=4). Chunk took 0.0 seconds.
Min u_count not hit, continuing.
Running chunk 3/10
Count of 110 for level Levenshtein distance <= 1 (cvv=4). Chunk took 0.1 seconds.
Exiting early since min count of 110 exceeds min_count_per_level = 100
Estimating u for: city (Comparison 4 of 5)
Running probe chunk (~1.00% of max_pairs)
Min u_count: 38 for comparison level Exact match on city (cvv=1)
Probe did not converge; restarting with normal chunking
Running chunk 1/10
Count of 928 for level Exact match on city (cvv=1). Chunk took 0.0 seconds.
Exiting early since min count of 928 exceeds min_count_per_level = 100
Estimating u for: email (Comparison 5 of 5)
Running probe chunk (~1.00% of max_pairs)
Min u_count: 1 for comparison level Jaro-Winkler distance of email >= 0.88 (cvv=2)
Probe did not converge; restarting with normal chunking
Running chunk 1/10
Count of 9 for level Jaro-Winkler >0.88 on username (cvv=1). Chunk took 0.1 seconds.
Min u_count not hit, continuing.
Running chunk 2/10
Count of 22 for level Jaro-Winkler >0.88 on username (cvv=1). Chunk took 0.1 seconds.
Min u_count not hit, continuing.
Running chunk 3/10
Count of 30 for level Jaro-Winkler >0.88 on username (cvv=1). Chunk took 0.1 seconds.
Min u_count not hit, continuing.
Running chunk 4/10
Count of 37 for level Jaro-Winkler >0.88 on username (cvv=1). Chunk took 0.1 seconds.
Min u_count not hit, continuing.
Running chunk 5/10
Count of 45 for level Jaro-Winkler >0.88 on username (cvv=1). Chunk took 0.1 seconds.
Min u_count not hit, continuing.
Running chunk 6/10
Count of 50 for level Jaro-Winkler >0.88 on username (cvv=1). Chunk took 0.1 seconds.
Min u_count not hit, continuing.
Running chunk 7/10
Count of 62 for level Jaro-Winkler >0.88 on username (cvv=1). Chunk took 0.1 seconds.
Min u_count not hit, continuing.
Running chunk 8/10
Count of 68 for level Jaro-Winkler >0.88 on username (cvv=1). Chunk took 0.1 seconds.
Min u_count not hit, continuing.
Running chunk 9/10
Count of 75 for level Jaro-Winkler >0.88 on username (cvv=1). Chunk took 0.1 seconds.
Min u_count not hit, continuing.
Running chunk 10/10
Count of 83 for level Jaro-Winkler >0.88 on username (cvv=1). Chunk took 0.1 seconds.
Min u_count not hit, continuing.
Estimated u probabilities using random sampling
Your model is not yet fully trained. Missing estimates for:
- first_name (no m values are trained).
- surname (no m values are trained).
- dob (no m values are trained).
- city (no m values are trained).
- email (no m values are trained).
session_dob = linker.training.estimate_parameters_using_expectation_maximisation(block_on("dob"))
session_email = linker.training.estimate_parameters_using_expectation_maximisation(
block_on("email")
)
session_first_name = linker.training.estimate_parameters_using_expectation_maximisation(
block_on("first_name")
)
----- Starting EM training session -----
[EM sampling] max_pairs is None — no sampling will be applied
Estimating the m probabilities of the model by blocking on:
l."dob" = r."dob"
Parameter estimates will be made for the following comparison(s):
- first_name
- surname
- city
- email
Parameter estimates cannot be made for the following comparison(s) since they are used in the blocking rules:
- dob
WARNING:
Level Jaro-Winkler >0.88 on username on comparison email not observed in dataset, unable to train m value
Iteration 1: Largest change in params was -0.412 in the m_probability of surname, level `Exact match on surname`
Iteration 2: Largest change in params was 0.125 in probability_two_random_records_match
Iteration 3: Largest change in params was 0.045 in the m_probability of first_name, level `All other comparisons`
Iteration 4: Largest change in params was 0.0156 in probability_two_random_records_match
Iteration 5: Largest change in params was 0.00652 in probability_two_random_records_match
Iteration 6: Largest change in params was 0.00314 in probability_two_random_records_match
Iteration 7: Largest change in params was 0.00192 in probability_two_random_records_match
Iteration 8: Largest change in params was 0.00187 in the m_probability of first_name, level `All other comparisons`
Iteration 9: Largest change in params was 0.00151 in the m_probability of first_name, level `All other comparisons`
Iteration 10: Largest change in params was 0.000926 in probability_two_random_records_match
Iteration 11: Largest change in params was 0.000548 in probability_two_random_records_match
Iteration 12: Largest change in params was 0.000319 in probability_two_random_records_match
Iteration 13: Largest change in params was 0.000185 in probability_two_random_records_match
Iteration 14: Largest change in params was 0.000107 in probability_two_random_records_match
Iteration 15: Largest change in params was 6.24e-05 in probability_two_random_records_match
EM converged after 15 iterations
m probability not trained for email - Jaro-Winkler >0.88 on username (comparison vector value: 1). This usually means the comparison level was never observed in the training data.
Your model is not yet fully trained. Missing estimates for:
- dob (no m values are trained).
- email (some m values are not trained).
----- Starting EM training session -----
[EM sampling] max_pairs is None — no sampling will be applied
Estimating the m probabilities of the model by blocking on:
l."email" = r."email"
Parameter estimates will be made for the following comparison(s):
- first_name
- surname
- dob
- city
Parameter estimates cannot be made for the following comparison(s) since they are used in the blocking rules:
- email
Iteration 1: Largest change in params was -0.439 in the m_probability of dob, level `Exact match on date of birth`
Iteration 2: Largest change in params was 0.0811 in probability_two_random_records_match
Iteration 3: Largest change in params was 0.0178 in probability_two_random_records_match
Iteration 4: Largest change in params was 0.00698 in probability_two_random_records_match
Iteration 5: Largest change in params was 0.00332 in probability_two_random_records_match
Iteration 6: Largest change in params was 0.00172 in probability_two_random_records_match
Iteration 7: Largest change in params was 0.000938 in probability_two_random_records_match
Iteration 8: Largest change in params was 0.000525 in probability_two_random_records_match
Iteration 9: Largest change in params was 0.000298 in probability_two_random_records_match
Iteration 10: Largest change in params was 0.000171 in probability_two_random_records_match
Iteration 11: Largest change in params was 9.83e-05 in probability_two_random_records_match
EM converged after 11 iterations
Your model is not yet fully trained. Missing estimates for:
- email (some m values are not trained).
----- Starting EM training session -----
[EM sampling] max_pairs is None — no sampling will be applied
Estimating the m probabilities of the model by blocking on:
l."first_name" = r."first_name"
Parameter estimates will be made for the following comparison(s):
- surname
- dob
- city
- email
Parameter estimates cannot be made for the following comparison(s) since they are used in the blocking rules:
- first_name
Iteration 1: Largest change in params was -0.182 in the m_probability of surname, level `All other comparisons`
Iteration 2: Largest change in params was -0.0144 in the m_probability of surname, level `All other comparisons`
Iteration 3: Largest change in params was -0.00402 in the m_probability of surname, level `All other comparisons`
Iteration 4: Largest change in params was -0.00183 in the m_probability of surname, level `All other comparisons`
Iteration 5: Largest change in params was -0.00106 in the m_probability of surname, level `All other comparisons`
Iteration 6: Largest change in params was -0.000679 in the m_probability of surname, level `All other comparisons`
Iteration 7: Largest change in params was -0.000452 in the m_probability of surname, level `All other comparisons`
Iteration 8: Largest change in params was -0.000306 in the m_probability of surname, level `All other comparisons`
Iteration 9: Largest change in params was -0.000208 in the m_probability of surname, level `All other comparisons`
Iteration 10: Largest change in params was -0.000142 in the m_probability of surname, level `All other comparisons`
Iteration 11: Largest change in params was -9.73e-05 in the m_probability of surname, level `All other comparisons`
EM converged after 11 iterations
Your model is fully trained. All comparisons have at least one estimate for their m and u values
results = linker.inference.predict(threshold_match_probability=0.9)
Blocking time: 0.02 seconds
Predict time (post-blocking): 0.08 seconds
results.as_duckdbpyrelation().limit(5).show(max_width=10000)
┌────────────────────┬────────────────────┬──────────────────┬──────────────────┬─────────────┬─────────────┬──────────────┬──────────────┬──────────────────┬───────────┬───────────┬───────────────┬────────────┬────────────┬───────────┬────────────┬────────────┬────────────┬───────────────────────┬───────────────────────┬─────────────┬───────────┐
│ match_weight │ match_probability │ source_dataset_l │ source_dataset_r │ unique_id_l │ unique_id_r │ first_name_l │ first_name_r │ gamma_first_name │ surname_l │ surname_r │ gamma_surname │ dob_l │ dob_r │ gamma_dob │ city_l │ city_r │ gamma_city │ email_l │ email_r │ gamma_email │ match_key │
│ double │ double │ varchar │ varchar │ int64 │ int64 │ varchar │ varchar │ int32 │ varchar │ varchar │ int32 │ date │ date │ int32 │ varchar │ varchar │ int32 │ varchar │ varchar │ int32 │ varchar │
├────────────────────┼────────────────────┼──────────────────┼──────────────────┼─────────────┼─────────────┼──────────────┼──────────────┼──────────────────┼───────────┼───────────┼───────────────┼────────────┼────────────┼───────────┼────────────┼────────────┼────────────┼───────────────────────┼───────────────────────┼─────────────┼───────────┤
│ 6.373348335504694 │ 0.9880814418247867 │ df_left │ df_right │ 47 │ 50 │ Erin │ Erin │ 4 │ Rogers │ Roers │ 3 │ 2010-01-02 │ 2010-03-03 │ 2 │ London │ NULL │ -1 │ e.rogers3@hopkins.org │ NULL │ -1 │ 0 │
│ 10.208614274517489 │ 0.9991556276020533 │ df_left │ df_right │ 69 │ 68 │ Reggie │ Reggie │ 4 │ Mason │ NULL │ -1 │ 1986-09-09 │ 1996-08-05 │ 1 │ Coventry │ Coventry │ 1 │ r..m51@leonard.net │ r.m51@leonard.net │ 2 │ 0 │
│ 9.744461232398068 │ 0.9988355552936555 │ df_left │ df_right │ 70 │ 68 │ Reggie │ Reggie │ 4 │ NULL │ NULL │ -1 │ 1986-09-09 │ 1996-08-05 │ 1 │ Coventry │ Coventry │ 1 │ r.m51@leonard.net │ r.m51@leonard.net │ 4 │ 0 │
│ 5.3479964748217 │ 0.9760360447073195 │ df_left │ df_right │ 71 │ 68 │ Reggie │ Reggie │ 4 │ Mason │ NULL │ -1 │ 1986-09-09 │ 1996-08-05 │ 1 │ NULL │ Coventry │ -1 │ r.m51@leonard.net │ r.m51@leonard.net │ 4 │ 0 │
│ 17.786165606563372 │ 0.9999955758614097 │ df_left │ df_right │ 89 │ 88 │ Lexi │ Lexi │ 4 │ NULL │ NULL │ -1 │ 1994-09-02 │ 1994-09-02 │ 5 │ Birmingham │ Birmingham │ 1 │ l.gordon34@french.com │ l.gordon34cfren@h.com │ 2 │ 0 │
└────────────────────┴────────────────────┴──────────────────┴──────────────────┴─────────────┴─────────────┴──────────────┴──────────────┴──────────────────┴───────────┴───────────┴───────────────┴────────────┴────────────┴───────────┴────────────┴────────────┴────────────┴───────────────────────┴───────────────────────┴─────────────┴───────────┘