Skip to content

Febrl4 link-only

Linking the febrl4 datasets

See A.2 here and here for the source of this data.

It consists of two datasets, A and B, of 5000 records each, with each record in dataset A having a corresponding record in dataset B. The aim will be to capture as many of those 5000 true links as possible, with minimal false linkages.

It is worth noting that we should not necessarily expect to capture all links. There are some links that although we know they do correspond to the same person, the data is so mismatched between them that we would not reasonably expect a model to link them, and indeed should a model do so may indicate that we have overengineered things using our knowledge of true links, which will not be a helpful reference in situations where we attempt to link unlabelled data, as will usually be the case.

Open In Colab

Exploring data and defining model

Firstly let's read in the data and have a little look at it

from splink import splink_datasets
from splink.internals.misc import show
import duckdb

df_a = splink_datasets.febrl4a
df_b = splink_datasets.febrl4b


def prepare_data(data):
    data = data.rename_columns([column.strip() for column in data.column_names])
    return (
        duckdb.sql(
            """
            select
                * replace (
                    trim(cast(date_of_birth as varchar)) as date_of_birth,
                    trim(cast(soc_sec_id as varchar)) as soc_sec_id,
                    trim(cast(postcode as varchar)) as postcode
                ),
                regexp_extract(rec_id, '^([^-]+-[^-]+)', 1) as cluster
            from data
            """
        )
        .arrow()
        .read_all()
    )


dfs = [prepare_data(dataset) for dataset in [df_a, df_b]]

show(dfs[0], rows=2)
show(dfs[1], rows=2)
downloading: https://raw.githubusercontent.com/moj-analytical-services/splink_datasets/master/data/febrl/dataset4a.csv



downloading: https://raw.githubusercontent.com/moj-analytical-services/splink_datasets/master/data/febrl/dataset4b.csv



┌──────────────┬────────────┬──────────┬───────────────┬────────────────────┬─────────────┬────────────────┬──────────┬─────────┬───────────────┬────────────┬──────────┐
│    rec_id    │ given_name │ surname  │ street_number │     address_1      │  address_2  │     suburb     │ postcode │  state  │ date_of_birth │ soc_sec_id │ cluster  │
│   varchar    │  varchar   │ varchar  │    varchar    │      varchar       │   varchar   │    varchar     │ varchar  │ varchar │    varchar    │  varchar   │ varchar  │
├──────────────┼────────────┼──────────┼───────────────┼────────────────────┼─────────────┼────────────────┼──────────┼─────────┼───────────────┼────────────┼──────────┤
│ rec-1070-org │  michaela  │  neumann │  8            │  stanley street    │  miami      │  winston hills │ 4223     │  nsw    │ 19151111      │ 5304218    │ rec-1070 │
│ rec-1016-org │  courtney  │  painter │  12           │  pinkerton circuit │  bega flats │  richlands     │ 4560     │  vic    │ 19161214      │ 4066625    │ rec-1016 │
└──────────────┴────────────┴──────────┴───────────────┴────────────────────┴─────────────┴────────────────┴──────────┴─────────┴───────────────┴────────────┴──────────┘

┌────────────────┬────────────┬─────────┬───────────────┬────────────────┬────────────┬─────────────┬──────────┬─────────┬───────────────┬────────────┬──────────┐
│     rec_id     │ given_name │ surname │ street_number │   address_1    │ address_2  │   suburb    │ postcode │  state  │ date_of_birth │ soc_sec_id │ cluster  │
│    varchar     │  varchar   │ varchar │    varchar    │    varchar     │  varchar   │   varchar   │ varchar  │ varchar │    varchar    │  varchar   │ varchar  │
├────────────────┼────────────┼─────────┼───────────────┼────────────────┼────────────┼─────────────┼──────────┼─────────┼───────────────┼────────────┼──────────┤
│ rec-561-dup-0  │  elton     │         │  3            │  light setreet │  pinehill  │  windermere │ 3212     │  vic    │ 19651013      │ 1551941    │ rec-561  │
│ rec-2642-dup-0 │  mitchell  │  maxon  │  47           │  edkins street │  lochaoair │  north ryde │ 3355     │  nsw    │ 19390212      │ 8859999    │ rec-2642 │
└────────────────┴────────────┴─────────┴───────────────┴────────────────┴────────────┴─────────────┴──────────┴─────────┴───────────────┴────────────┴──────────┘

Next, to better understand which variables will prove useful in linking, we have a look at how populated each column is, as well as the distribution of unique values within each

from splink import DuckDBAPI, Linker, SettingsCreator

basic_settings = SettingsCreator(
    unique_id_column_name="rec_id",
    link_type="link_only",
    # NB as we are linking one-one, we know the probability that a random pair will be a match
    # hence we could set:
    # "probability_two_random_records_match": 1/5000,
    # however we will not specify this here, as we will use this as a check that
    # our estimation procedure returns something sensible
)

db_api = DuckDBAPI()
dfs_sdf = [db_api.register(df) for df in dfs]
linker = Linker(dfs_sdf, basic_settings)

It's usually a good idea to perform exploratory analysis on your data so you understand what's in each column and how often it's missing

from splink.exploratory import completeness_chart

db_api = DuckDBAPI()
dfs_sdf = [db_api.register(df) for df in dfs]
completeness_chart(dfs_sdf)
from splink.exploratory import profile_columns

db_api = DuckDBAPI()
dfs_sdf = [db_api.register(df) for df in dfs]
profile_columns(dfs_sdf, column_expressions=["given_name", "surname"])

Next let's come up with some candidate blocking rules, which define which record comparisons are generated, and have a look at how many comparisons each will generate.

For blocking rules that we use in prediction, our aim is to have the union of all rules cover all true matches, whilst avoiding generating so many comparisons that it becomes computationally intractable - i.e. each true match should have at least one of the following conditions holding.

from splink import DuckDBAPI, block_on
from splink.blocking_analysis import (
    chart_comparisons_from_blocking_rules,
)

blocking_rules = [
    block_on("given_name", "surname"),
    # A blocking rule can also be an aribtrary SQL expression
    "l.given_name = r.surname and l.surname = r.given_name",
    block_on("date_of_birth"),
    block_on("soc_sec_id"),
    block_on("state", "address_1"),
    block_on("street_number", "address_1"),
    block_on("postcode"),
]


db_api = DuckDBAPI()
dfs_sdf = [db_api.register(df) for df in dfs]
chart_comparisons_from_blocking_rules(
    dfs_sdf,
    blocking_rules=blocking_rules,
    link_type="link_only",
    unique_id_column_name="rec_id",
    source_dataset_column_name="source_dataset",
    record_sample_proportion=1.0,
)

The broadest rule, having a matching postcode, unsurpisingly gives the largest number of comparisons. For this small dataset we still have a very manageable number, but if it was larger we might have needed to include a further AND condition with it to break the number of comparisons further.

Now we get the full settings by including the blocking rules, as well as deciding the actual comparisons we will be including in our model.

We will define two models, each with a separate linker with different settings, so that we can compare performance. One will be a very basic model, whilst the other will include a lot more detail.

import splink.comparison_level_library as cll
import splink.comparison_library as cl


# the simple model only considers a few columns, and only two comparison levels for each
simple_model_settings = SettingsCreator(
    unique_id_column_name="rec_id",
    link_type="link_only",
    blocking_rules_to_generate_predictions=blocking_rules,
    comparisons=[
        cl.ExactMatch("given_name").configure(term_frequency_adjustments=True),
        cl.ExactMatch("surname").configure(term_frequency_adjustments=True),
        cl.ExactMatch("street_number").configure(term_frequency_adjustments=True),
    ],
    retain_intermediate_calculation_columns=True,
)

# the detailed model considers more columns, using the information we saw in the exploratory phase
# we also include further comparison levels to account for typos and other differences
detailed_model_settings = SettingsCreator(
    unique_id_column_name="rec_id",
    link_type="link_only",
    blocking_rules_to_generate_predictions=blocking_rules,
    comparisons=[
        cl.NameComparison("given_name").configure(term_frequency_adjustments=True),
        cl.NameComparison("surname").configure(term_frequency_adjustments=True),
        cl.DateOfBirthComparison(
            "date_of_birth",
            input_is_string=True,
            datetime_format="%Y%m%d",
            invalid_dates_as_null=True,
        ),
        cl.DamerauLevenshteinAtThresholds("soc_sec_id", [1, 2]),
        cl.ExactMatch("street_number").configure(term_frequency_adjustments=True),
        cl.DamerauLevenshteinAtThresholds("postcode", [1, 2]).configure(
            term_frequency_adjustments=True
        ),
        # we don't consider further location columns as they will be strongly correlated with postcode
    ],
    retain_intermediate_calculation_columns=True,
)


db_api = DuckDBAPI()
dfs_sdf = [db_api.register(df) for df in dfs]
linker_simple = Linker(dfs_sdf, simple_model_settings)
linker_detailed = Linker(dfs_sdf, detailed_model_settings)

Estimating model parameters

We need to furnish our models with parameter estimates so that we can generate results. We will focus on the detailed model, generating the values for the simple model at the end

We can instead estimate the probability two random records match, and compare with the known value of 1/5000 = 0.0002, to see how well our estimation procedure works.

To do this we come up with some deterministic rules - the aim here is that we generate very few false positives (i.e. we expect that the majority of records with at least one of these conditions holding are true matches), whilst also capturing the majority of matches - our guess here is that these two rules should capture 80% of all matches.

deterministic_rules = [
    block_on("soc_sec_id"),
    block_on("given_name", "surname", "date_of_birth"),
]

linker_detailed.training.estimate_probability_two_random_records_match(
    deterministic_rules, recall=0.8
)
Probability two random records match is estimated to be  0.000239.
This means that amongst all possible pairwise record comparisons, one in 4,185.85 are expected to match.  With 25,000,000 total possible comparisons, we expect a total of around 5,972.50 matching pairs

Even playing around with changing these deterministic rules, or the nominal recall leaves us with an answer which is pretty close to our known value

Next we estimate u and m values for each comparison, so that we can move to generating predictions

# We generally recommend setting max pairs higher (e.g. 1e7 or more)
# But this will run faster for the purpose of this demo
linker_detailed.training.estimate_u_using_random_sampling(max_pairs=1e6)
You are using the default value for `max_pairs`, which may be too small and thus lead to inaccurate estimates for your model's u-parameters. Consider increasing to 1e8 or 1e9, which will result in more accurate estimates, but with a longer run time.


----- Estimating u probabilities using random sampling -----


Estimating u with: max_pairs = 1,000,000, min_count_per_level = 100, num_chunks = 10



Estimating u for: given_name (Comparison 1 of 6)


  Running probe chunk (~1.00% of max_pairs)


  Min u_count: 19 for comparison level Jaro-Winkler distance of given_name >= 0.88 (cvv=2)


  Probe did not converge; restarting with normal chunking



  Running chunk 1/10


  Count of 183 for level Jaro-Winkler distance of given_name >= 0.88 (cvv=2). Chunk took 0.1 seconds.


  Exiting early since min count of 183 exceeds min_count_per_level = 100



Estimating u for: surname (Comparison 2 of 6)


  Running probe chunk (~1.00% of max_pairs)


  Min u_count: 6 for comparison level Jaro-Winkler distance of surname >= 0.88 (cvv=2)


  Probe did not converge; restarting with normal chunking



  Running chunk 1/10


  Count of 81 for level Jaro-Winkler distance of surname >= 0.88 (cvv=2). Chunk took 0.1 seconds.


  Min u_count not hit, continuing.


  Running chunk 2/10


  Count of 164 for level Jaro-Winkler distance of surname >= 0.88 (cvv=2). Chunk took 0.1 seconds.


  Exiting early since min count of 164 exceeds min_count_per_level = 100



Estimating u for: date_of_birth (Comparison 3 of 6)


  Running probe chunk (~1.00% of max_pairs)


  Min u_count: 2 for comparison level Exact match on date of birth (cvv=5)


  Probe did not converge; restarting with normal chunking



  Running chunk 1/10


  Count of 24 for level Exact match on date of birth (cvv=5). Chunk took 0.1 seconds.


  Min u_count not hit, continuing.


  Running chunk 2/10


  Count of 44 for level Exact match on date of birth (cvv=5). Chunk took 0.2 seconds.


  Min u_count not hit, continuing.


  Running chunk 3/10


  Count of 59 for level Exact match on date of birth (cvv=5). Chunk took 0.2 seconds.


  Min u_count not hit, continuing.


  Running chunk 4/10


  Count of 76 for level Exact match on date of birth (cvv=5). Chunk took 0.2 seconds.


  Min u_count not hit, continuing.


  Running chunk 5/10


  Count of 97 for level Exact match on date of birth (cvv=5). Chunk took 0.3 seconds.


  Min u_count not hit, continuing.


  Running chunk 6/10


  Count of 114 for level Exact match on date of birth (cvv=5). Chunk took 0.3 seconds.


  Exiting early since min count of 114 exceeds min_count_per_level = 100



Estimating u for: soc_sec_id (Comparison 4 of 6)


  Running probe chunk (~1.00% of max_pairs)


  Min u_count: 0 for comparison level Damerau-Levenshtein distance of soc_sec_id <= 1 (cvv=2)


  Probe did not converge; restarting with normal chunking



  Running chunk 1/10


  Count of 1 for level Damerau-Levenshtein distance of soc_sec_id <= 1 (cvv=2). Chunk took 0.3 seconds.


  Min u_count not hit, continuing.


  Running chunk 2/10


  Count of 2 for level Damerau-Levenshtein distance of soc_sec_id <= 1 (cvv=2). Chunk took 0.4 seconds.


  Min u_count not hit, continuing.


  Running chunk 3/10


  Count of 4 for level Damerau-Levenshtein distance of soc_sec_id <= 1 (cvv=2). Chunk took 0.3 seconds.


  Min u_count not hit, continuing.


  Running chunk 4/10


  Count of 8 for level Damerau-Levenshtein distance of soc_sec_id <= 1 (cvv=2). Chunk took 0.4 seconds.


  Min u_count not hit, continuing.


  Running chunk 5/10


  Count of 8 for level Damerau-Levenshtein distance of soc_sec_id <= 1 (cvv=2). Chunk took 0.3 seconds.


  Min u_count not hit, continuing.


  Running chunk 6/10


  Count of 10 for level Damerau-Levenshtein distance of soc_sec_id <= 1 (cvv=2). Chunk took 0.3 seconds.


  Min u_count not hit, continuing.


  Running chunk 7/10


  Count of 11 for level Damerau-Levenshtein distance of soc_sec_id <= 1 (cvv=2). Chunk took 0.3 seconds.


  Min u_count not hit, continuing.


  Running chunk 8/10


  Count of 12 for level Damerau-Levenshtein distance of soc_sec_id <= 1 (cvv=2). Chunk took 0.6 seconds.


  Min u_count not hit, continuing.


  Running chunk 9/10


  Count of 17 for level Damerau-Levenshtein distance of soc_sec_id <= 1 (cvv=2). Chunk took 0.4 seconds.


  Min u_count not hit, continuing.


  Running chunk 10/10


  Count of 18 for level Damerau-Levenshtein distance of soc_sec_id <= 1 (cvv=2). Chunk took 0.4 seconds.


  Min u_count not hit, continuing.



Estimating u for: street_number (Comparison 5 of 6)


  Running probe chunk (~1.00% of max_pairs)


  Min u_count: 233 for comparison level Exact match on street_number (cvv=1)


  Exiting early since min count of 233 exceeds min_count_per_level = 100



Estimating u for: postcode (Comparison 6 of 6)


  Running probe chunk (~1.00% of max_pairs)


  Min u_count: 21 for comparison level Exact match on postcode (cvv=3)


  Probe did not converge; restarting with normal chunking



  Running chunk 1/10


  Count of 121 for level Exact match on postcode (cvv=3). Chunk took 0.4 seconds.


  Exiting early since min count of 121 exceeds min_count_per_level = 100



Estimated u probabilities using random sampling



Your model is not yet fully trained. Missing estimates for:
    - given_name (no m values are trained).
    - surname (no m values are trained).
    - date_of_birth (no m values are trained).
    - soc_sec_id (no m values are trained).
    - street_number (no m values are trained).
    - postcode (no m values are trained).

When training the m values using expectation maximisation, we need somre more blocking rules to reduce the total number of comparisons. For each rule, we want to ensure that we have neither proportionally too many matches, or too few.

We must run this multiple times using different rules so that we can obtain estimates for all comparisons - if we block on e.g. date_of_birth, then we cannot compute the m values for the date_of_birth comparison, as we have only looked at records where these match.

session_dob = (
    linker_detailed.training.estimate_parameters_using_expectation_maximisation(
        block_on("date_of_birth"), estimate_without_term_frequencies=True
    )
)
session_pc = (
    linker_detailed.training.estimate_parameters_using_expectation_maximisation(
        block_on("postcode"), estimate_without_term_frequencies=True
    )
)
----- Starting EM training session -----



[EM sampling] max_pairs is None — no sampling will be applied


Estimating the m probabilities of the model by blocking on:
l."date_of_birth" = r."date_of_birth"

Parameter estimates will be made for the following comparison(s):
    - given_name
    - surname
    - soc_sec_id
    - street_number
    - postcode

Parameter estimates cannot be made for the following comparison(s) since they are used in the blocking rules: 
    - date_of_birth





Iteration 1: Largest change in params was -0.353 in probability_two_random_records_match


Iteration 2: Largest change in params was 0.00355 in the m_probability of given_name, level `All other comparisons`


Iteration 3: Largest change in params was 9.73e-05 in the m_probability of soc_sec_id, level `All other comparisons`



EM converged after 3 iterations



Your model is not yet fully trained. Missing estimates for:
    - date_of_birth (no m values are trained).



----- Starting EM training session -----



[EM sampling] max_pairs is None — no sampling will be applied


Estimating the m probabilities of the model by blocking on:
l."postcode" = r."postcode"

Parameter estimates will be made for the following comparison(s):
    - given_name
    - surname
    - date_of_birth
    - soc_sec_id
    - street_number

Parameter estimates cannot be made for the following comparison(s) since they are used in the blocking rules: 
    - postcode





Iteration 1: Largest change in params was 0.0284 in the m_probability of date_of_birth, level `All other comparisons`


Iteration 2: Largest change in params was 0.00039 in the m_probability of date_of_birth, level `All other comparisons`


Iteration 3: Largest change in params was 8.92e-06 in the m_probability of soc_sec_id, level `All other comparisons`



EM converged after 3 iterations



Your model is fully trained. All comparisons have at least one estimate for their m and u values

If we wish we can have a look at how our parameter estimates changes over these training sessions

session_dob.m_u_values_interactive_history_chart()

For variables that aren't used in the m-training blocking rules, we have two estimates --- one from each of the training sessions (see for example street_number). We can have a look at how the values compare between them, to ensure that we don't have drastically different values, which may be indicative of an issue.

linker_detailed.visualisations.parameter_estimate_comparisons_chart()

We repeat our parameter estimations for the simple model in much the same fashion

linker_simple.training.estimate_probability_two_random_records_match(
    deterministic_rules, recall=0.8
)
linker_simple.training.estimate_u_using_random_sampling(max_pairs=1e7)
session_ssid = (
    linker_simple.training.estimate_parameters_using_expectation_maximisation(
        block_on("given_name"), estimate_without_term_frequencies=True
    )
)
session_pc = linker_simple.training.estimate_parameters_using_expectation_maximisation(
    block_on("street_number"), estimate_without_term_frequencies=True
)
linker_simple.visualisations.parameter_estimate_comparisons_chart()
Probability two random records match is estimated to be  0.000239.
This means that amongst all possible pairwise record comparisons, one in 4,185.85 are expected to match.  With 25,000,000 total possible comparisons, we expect a total of around 5,972.50 matching pairs


----- Estimating u probabilities using random sampling -----


Estimating u with: max_pairs = 10,000,000, min_count_per_level = 100, num_chunks = 10



Estimating u for: given_name (Comparison 1 of 3)


  Running probe chunk (~1.00% of max_pairs)


  Min u_count: 656 for comparison level Exact match on given_name (cvv=1)


  Exiting early since min count of 656 exceeds min_count_per_level = 100



Estimating u for: surname (Comparison 2 of 3)


  Running probe chunk (~1.00% of max_pairs)


  Min u_count: 612 for comparison level Exact match on surname (cvv=1)


  Exiting early since min count of 612 exceeds min_count_per_level = 100



Estimating u for: street_number (Comparison 3 of 3)


  Running probe chunk (~1.00% of max_pairs)


  Min u_count: 2,020 for comparison level Exact match on street_number (cvv=1)


  Exiting early since min count of 2,020 exceeds min_count_per_level = 100



Estimated u probabilities using random sampling



Your model is not yet fully trained. Missing estimates for:
    - given_name (no m values are trained).
    - surname (no m values are trained).
    - street_number (no m values are trained).



----- Starting EM training session -----



[EM sampling] max_pairs is None — no sampling will be applied


Estimating the m probabilities of the model by blocking on:
l."given_name" = r."given_name"

Parameter estimates will be made for the following comparison(s):
    - surname
    - street_number

Parameter estimates cannot be made for the following comparison(s) since they are used in the blocking rules: 
    - given_name





Iteration 1: Largest change in params was 0.0721 in the m_probability of surname, level `All other comparisons`


Iteration 2: Largest change in params was -0.0272 in the m_probability of surname, level `Exact match on surname`


Iteration 3: Largest change in params was -0.0261 in the m_probability of surname, level `Exact match on surname`


Iteration 4: Largest change in params was 0.0241 in the m_probability of surname, level `All other comparisons`


Iteration 5: Largest change in params was -0.0209 in the m_probability of surname, level `Exact match on surname`


Iteration 6: Largest change in params was 0.0173 in the m_probability of surname, level `All other comparisons`


Iteration 7: Largest change in params was -0.0137 in the m_probability of surname, level `Exact match on surname`


Iteration 8: Largest change in params was -0.0105 in the m_probability of surname, level `Exact match on surname`


Iteration 9: Largest change in params was 0.00791 in the m_probability of surname, level `All other comparisons`


Iteration 10: Largest change in params was -0.00586 in the m_probability of surname, level `Exact match on surname`


Iteration 11: Largest change in params was -0.0043 in the m_probability of surname, level `Exact match on surname`


Iteration 12: Largest change in params was -0.00315 in the m_probability of surname, level `Exact match on surname`


Iteration 13: Largest change in params was -0.0023 in the m_probability of surname, level `Exact match on surname`


Iteration 14: Largest change in params was -0.00169 in the m_probability of surname, level `Exact match on surname`


Iteration 15: Largest change in params was -0.00124 in the m_probability of surname, level `Exact match on surname`


Iteration 16: Largest change in params was -0.000911 in the m_probability of surname, level `Exact match on surname`


Iteration 17: Largest change in params was 0.000673 in the m_probability of surname, level `All other comparisons`


Iteration 18: Largest change in params was -0.000498 in the m_probability of surname, level `Exact match on surname`


Iteration 19: Largest change in params was -0.00037 in the m_probability of surname, level `Exact match on surname`


Iteration 20: Largest change in params was -0.000275 in the m_probability of surname, level `Exact match on surname`


Iteration 21: Largest change in params was -0.000206 in the m_probability of surname, level `Exact match on surname`


Iteration 22: Largest change in params was -0.000154 in the m_probability of surname, level `Exact match on surname`


Iteration 23: Largest change in params was -0.000115 in the m_probability of surname, level `Exact match on surname`


Iteration 24: Largest change in params was 8.63e-05 in the m_probability of surname, level `All other comparisons`



EM converged after 24 iterations



Your model is not yet fully trained. Missing estimates for:
    - given_name (no m values are trained).



----- Starting EM training session -----



[EM sampling] max_pairs is None — no sampling will be applied


Estimating the m probabilities of the model by blocking on:
l."street_number" = r."street_number"

Parameter estimates will be made for the following comparison(s):
    - given_name
    - surname

Parameter estimates cannot be made for the following comparison(s) since they are used in the blocking rules: 
    - street_number





Iteration 1: Largest change in params was -0.0397 in the m_probability of surname, level `Exact match on surname`


Iteration 2: Largest change in params was 0.0365 in the m_probability of surname, level `Exact match on surname`


Iteration 3: Largest change in params was -0.025 in the m_probability of surname, level `All other comparisons`


Iteration 4: Largest change in params was 0.017 in the m_probability of surname, level `Exact match on surname`


Iteration 5: Largest change in params was 0.0119 in the m_probability of surname, level `Exact match on surname`


Iteration 6: Largest change in params was -0.00865 in the m_probability of given_name, level `Exact match on given_name`


Iteration 7: Largest change in params was -0.00767 in the m_probability of given_name, level `Exact match on given_name`


Iteration 8: Largest change in params was -0.00672 in the m_probability of given_name, level `Exact match on given_name`


Iteration 9: Largest change in params was 0.00582 in the m_probability of given_name, level `All other comparisons`


Iteration 10: Largest change in params was 0.00498 in the m_probability of given_name, level `All other comparisons`


Iteration 11: Largest change in params was -0.00423 in the m_probability of given_name, level `Exact match on given_name`


Iteration 12: Largest change in params was 0.00357 in the m_probability of given_name, level `All other comparisons`


Iteration 13: Largest change in params was -0.00299 in the m_probability of given_name, level `Exact match on given_name`


Iteration 14: Largest change in params was 0.0025 in the m_probability of given_name, level `All other comparisons`


Iteration 15: Largest change in params was -0.00208 in the m_probability of given_name, level `Exact match on given_name`


Iteration 16: Largest change in params was 0.00172 in the m_probability of given_name, level `All other comparisons`


Iteration 17: Largest change in params was -0.00143 in the m_probability of given_name, level `Exact match on given_name`


Iteration 18: Largest change in params was -0.00118 in the m_probability of given_name, level `Exact match on given_name`


Iteration 19: Largest change in params was 0.000974 in the m_probability of given_name, level `All other comparisons`


Iteration 20: Largest change in params was -0.000805 in the m_probability of given_name, level `Exact match on given_name`


Iteration 21: Largest change in params was -0.000666 in the m_probability of given_name, level `Exact match on given_name`


Iteration 22: Largest change in params was 0.000551 in the m_probability of given_name, level `All other comparisons`


Iteration 23: Largest change in params was 0.000456 in the m_probability of given_name, level `All other comparisons`


Iteration 24: Largest change in params was -0.000378 in the m_probability of given_name, level `Exact match on given_name`


Iteration 25: Largest change in params was 0.000314 in the m_probability of given_name, level `All other comparisons`



EM converged after 25 iterations



Your model is fully trained. All comparisons have at least one estimate for their m and u values
# import json
# we can have a look at the full settings if we wish, including the values of our estimated parameters:
# print(json.dumps(linker_detailed._settings_obj.as_dict(), indent=2))
# we can also get a handy summary of of the model in an easily readable format if we wish:
# print(linker_detailed._settings_obj.human_readable_description)
# (we suppress output here for brevity)

We can now visualise some of the details of our models. We can look at the match weights, which tell us the relative importance for/against a match for each of our comparsion levels.

Comparing the two models will show the added benefit we get in the more detailed model --- what in the simple model is classed as 'all other comparisons' is instead broken down further, and we can see that the detail of how this is broken down in fact gives us quite a bit of useful information about the likelihood of a match.

linker_simple.visualisations.match_weights_chart()
linker_detailed.visualisations.match_weights_chart()

As well as the match weights, which give us an idea of the overall effect of each comparison level, we can also look at the individual u and m parameter estimates, which tells us about the prevalence of coincidences and mistakes (for further details/explanation about this see this article). We might want to revise aspects of our model based on the information we ascertain here.

Note however that some of these values are very small, which is why the match weight chart is often more useful for getting a decent picture of things.

# linker_simple.m_u_parameters_chart()
linker_detailed.visualisations.m_u_parameters_chart()

It is also useful to have a look at unlinkable records - these are records which do not contain enough information to be linked at some match probability threshold. We can figure this out be seeing whether records are able to be matched with themselves.

This is of course relative to the information we have put into the model - we see that in our simple model, at a 99% match threshold nearly 10% of records are unlinkable, as we have not included enough information in the model for distinct records to be adequately distinguished; this is not an issue in our more detailed model.

linker_simple.evaluation.unlinkables_chart()
linker_detailed.evaluation.unlinkables_chart()

Our simple model doesn't do terribly, but suffers if we want to have a high match probability --- to be 99% (match weight ~7) certain of matches we have ~10% of records that we will be unable to link.

Our detailed model, however, has enough nuance that we can at least self-link records.

Predictions

Now that we have had a look into the details of the models, we will focus on only our more detailed model, which should be able to capture more of the genuine links in our data

predictions = linker_detailed.inference.predict(threshold_match_probability=0.2)
predictions.as_duckdbpyrelation().limit(5).show(max_width=10000)
Blocking time: 0.21 seconds


Predict time (post-blocking): 0.60 seconds


┌───────────────────┬────────────────────┬─────────────────────────┬─────────────────────────┬──────────────┬────────────────┬──────────────┬──────────────┬──────────────────┬─────────────────┬─────────────────┬──────────────────┬──────────────────────┬───────────┬───────────┬───────────────┬──────────────┬──────────────┬───────────────────┬────────────────────┬─────────────────┬─────────────────┬─────────────────────┬────────────────────┬──────────────┬──────────────┬──────────────────┬────────────────────┬─────────────────┬─────────────────┬─────────────────────┬────────────────────┬────────────────────┬───────────────────┬─────────────────────────┬────────────┬────────────┬────────────────┬───────────────┬───────────────┬──────────────────┬──────────────────────┬─────────┬─────────┬──────────────────────┬────────────────────────────────────┬───────────┐
│   match_weight    │ match_probability  │    source_dataset_l     │    source_dataset_r     │   rec_id_l   │    rec_id_r    │ given_name_l │ given_name_r │ gamma_given_name │ tf_given_name_l │ tf_given_name_r │  mw_given_name   │ mw_tf_adj_given_name │ surname_l │ surname_r │ gamma_surname │ tf_surname_l │ tf_surname_r │    mw_surname     │ mw_tf_adj_surname  │ date_of_birth_l │ date_of_birth_r │ gamma_date_of_birth │  mw_date_of_birth  │ soc_sec_id_l │ soc_sec_id_r │ gamma_soc_sec_id │   mw_soc_sec_id    │ street_number_l │ street_number_r │ gamma_street_number │ tf_street_number_l │ tf_street_number_r │ mw_street_number  │ mw_tf_adj_street_number │ postcode_l │ postcode_r │ gamma_postcode │ tf_postcode_l │ tf_postcode_r │   mw_postcode    │  mw_tf_adj_postcode  │ state_l │ state_r │     address_1_l      │            address_1_r             │ match_key │
│      double       │       double       │         varchar         │         varchar         │   varchar    │    varchar     │   varchar    │   varchar    │      int32       │     double      │     double      │      double      │        double        │  varchar  │  varchar  │     int32     │    double    │    double    │      double       │       double       │     varchar     │     varchar     │        int32        │       double       │   varchar    │   varchar    │      int32       │       double       │     varchar     │     varchar     │        int32        │       double       │       double       │      double       │         double          │  varchar   │  varchar   │     int32      │    double     │    double     │      double      │        double        │ varchar │ varchar │       varchar        │              varchar               │  varchar  │
├───────────────────┼────────────────────┼─────────────────────────┼─────────────────────────┼──────────────┼────────────────┼──────────────┼──────────────┼──────────────────┼─────────────────┼─────────────────┼──────────────────┼──────────────────────┼───────────┼───────────┼───────────────┼──────────────┼──────────────┼───────────────────┼────────────────────┼─────────────────┼─────────────────┼─────────────────────┼────────────────────┼──────────────┼──────────────┼──────────────────┼────────────────────┼─────────────────┼─────────────────┼─────────────────────┼────────────────────┼────────────────────┼───────────────────┼─────────────────────────┼────────────┼────────────┼────────────────┼───────────────┼───────────────┼──────────────────┼──────────────────────┼─────────┼─────────┼──────────────────────┼────────────────────────────────────┼───────────┤
│ 43.45796162794336 │ 0.9999999999999172 │ __splink__input_table_0 │ __splink__input_table_1 │ rec-1016-org │ rec-1016-dup-0 │  courtney    │  courtney    │                4 │          0.0021 │          0.0021 │ 7.27625992308283 │   1.0545800636928844 │  painter  │  painter  │             4 │       0.0012 │       0.0012 │ 7.766001212885147 │ 1.3637920659558702 │ 19161214        │ 19161214        │                   5 │ 12.267594953187963 │ 4066625      │ 4066625      │                3 │ 12.356798809570432 │  12             │  12             │                   1 │             0.0218 │             0.0218 │ 5.807560458292876 │     -0.5358009350467752 │ 4560       │ 4560       │              3 │         0.003 │         0.003 │ 9.25446635484313 │  -1.1223304537872352 │  vic    │  vci    │  pinkerton circuit   │  pinkerton circuit                 │ 0         │
│  45.5272483993876 │ 0.9999999999999802 │ __splink__input_table_0 │ __splink__input_table_1 │ rec-4405-org │ rec-4405-dup-0 │  charles     │  charles     │                4 │          0.0017 │          0.0017 │ 7.27625992308283 │    1.359434645221305 │  green    │  green    │             4 │       0.0165 │       0.0165 │ 7.766001212885147 │ -2.417567647568789 │ 19480930        │ 19480930        │                   5 │ 12.267594953187963 │ 4365168      │ 4365168      │                3 │ 12.356798809570432 │  38             │  38             │                   1 │              0.007 │              0.007 │ 5.807560458292876 │      1.1031003727851854 │ 4566       │ 4566       │              3 │        0.0002 │        0.0002 │ 9.25446635484313 │   2.7845601418212826 │  nsw    │  nsw    │  salkauskas crescent │  salkauskas crescent               │ 0         │
│  52.0335008747779 │ 0.9999999999999998 │ __splink__input_table_0 │ __splink__input_table_1 │ rec-1288-org │ rec-1288-dup-0 │  vanessa     │  vanessa     │                4 │          0.0008 │          0.0008 │ 7.27625992308283 │    2.446897486471644 │  parr     │  parr     │             4 │       0.0018 │       0.0018 │ 7.766001212885147 │ 0.7788295652347141 │ 19951119        │ 19951119        │                   5 │ 12.267594953187963 │ 9239102      │ 9239102      │                3 │ 12.356798809570432 │  905            │  905            │                   1 │             0.0002 │             0.0002 │ 5.807560458292876 │       6.232383389730151 │ 2135       │ 2135       │              3 │        0.0015 │        0.0015 │ 9.25446635484313 │ -0.12233045378723517 │  sa     │  sa     │  macquoid place      │  macquoid place                    │ 0         │
│ 50.32895675830407 │ 0.9999999999999993 │ __splink__input_table_0 │ __splink__input_table_1 │ rec-3585-org │ rec-3585-dup-0 │  mikayla     │  mikayla     │                4 │          0.0011 │          0.0011 │ 7.27625992308283 │   1.9874658678343478 │  malloney │  malloney │             4 │       0.0002 │       0.0002 │ 7.766001212885147 │ 3.9487545666770263 │ 19860208        │ 19860208        │                   5 │ 12.267594953187963 │ 7207688      │ 7207688      │                3 │ 12.356798809570432 │  37             │  37             │                   1 │              0.008 │              0.008 │ 5.807560458292876 │      0.9104552948427891 │ 4552       │ 4552       │              3 │        0.0008 │        0.0008 │ 9.25446635484313 │   0.7845601418212826 │  vic    │  vic    │  randwick road       │  randwick road                     │ 0         │
│ 37.85053237804072 │  0.999999999995965 │ __splink__input_table_0 │ __splink__input_table_1 │ rec-298-org  │ rec-298-dup-0  │  blake       │  blake       │                4 │          0.0038 │          0.0038 │ 7.27625992308283 │  0.19896997302805897 │  howie    │  howie    │             4 │       0.0009 │       0.0009 │ 7.766001212885147 │ 1.7788295652347141 │ 19250301        │ 19250301        │                   5 │ 12.267594953187963 │ 5180548      │ 5180548      │                3 │ 12.356798809570432 │  1              │  1              │                   1 │             0.0332 │             0.0332 │ 5.807560458292876 │     -1.1426560416167737 │ 6017       │ 6071       │              2 │        0.0005 │        0.0005 │ 3.57213434910923 │                  0.0 │  vic    │  vic    │  cutlack street      │  belmont park belted galloway stud │ 0         │
└───────────────────┴────────────────────┴─────────────────────────┴─────────────────────────┴──────────────┴────────────────┴──────────────┴──────────────┴──────────────────┴─────────────────┴─────────────────┴──────────────────┴──────────────────────┴───────────┴───────────┴───────────────┴──────────────┴──────────────┴───────────────────┴────────────────────┴─────────────────┴─────────────────┴─────────────────────┴────────────────────┴──────────────┴──────────────┴──────────────────┴────────────────────┴─────────────────┴─────────────────┴─────────────────────┴────────────────────┴────────────────────┴───────────────────┴─────────────────────────┴────────────┴────────────┴────────────────┴───────────────┴───────────────┴──────────────────┴──────────────────────┴─────────┴─────────┴──────────────────────┴────────────────────────────────────┴───────────┘

We can see how our model performs at different probability thresholds, with a couple of options depending on the space we wish to view things

linker_detailed.evaluation.accuracy_analysis_from_labels_column(
    "cluster", output_type="accuracy"
)
Blocking time: 0.13 seconds


Predict time (post-blocking): 0.81 seconds

and we can easily see how many individuals we identify and link by looking at clusters generated at some threshold match probability of interest - in this example 99%

clusters = linker_detailed.clustering.cluster_pairwise_predictions_at_threshold(
    predictions, threshold_match_probability=0.99
)
cluster_size_distribution = clusters.query_sql(
    """
    select cluster_size, count(*) as cluster_count
    from (
        select cluster_id, count(*) as cluster_size
        from {this}
        group by cluster_id
    )
    group by cluster_size
    order by cluster_size
    """
)
cluster_size_distribution.as_duckdbpyrelation().show()
Completed iteration 1, num edges remaining to process: 0


┌──────────────┬───────────────┐
│ cluster_size │ cluster_count │
│    int64     │     int64     │
├──────────────┼───────────────┤
│            1 │            82 │
│            2 │          4959 │
└──────────────┴───────────────┘

In this case, we happen to know what the true links are, so we can manually inspect the ones that are doing worst to see what our model is not capturing - i.e. where we have false negatives.

Similarly, we can look at the non-links which are performing the best, to see whether we have an issue with false positives.

Ordinarily we would not have this luxury, and so would need to dig a bit deeper for clues as to how to improve our model, such as manually inspecting records across threshold probabilities,

df_predictions_with_clusters = predictions.query_sql(
    """
    select
        *,
        regexp_extract(rec_id_l, '^([^-]+-[^-]+)', 1) as cluster_l,
        regexp_extract(rec_id_r, '^([^-]+-[^-]+)', 1) as cluster_r
    from {this}
    """
)
df_true_links = df_predictions_with_clusters.query_sql(
    """
    select *
    from {this}
    where cluster_l = cluster_r
    order by match_probability
    """
)
records_to_view = 3
linker_detailed.visualisations.waterfall_chart(
    df_true_links.as_record_list(limit=records_to_view)
)
df_non_links = df_predictions_with_clusters.query_sql(
    """
    select *
    from {this}
    where cluster_l != cluster_r
    order by match_probability desc
    """
)
linker_detailed.visualisations.waterfall_chart(
    df_non_links.as_record_list(limit=records_to_view)
)

Further refinements

Looking at the non-links we have done well in having no false positives at any substantial match probability --- however looking at some of the true links we can see that there are a few that we are not capturing with sufficient match probability.

We can see that there are a few features that we are not capturing/weighting appropriately

  • single-character transpostions, particularly in postcode (which is being lumped in with more 'severe typos'/probable non-matches)
  • given/sur-names being swapped with typos
  • given/sur-names being cross-matches on one only, with no match on the other cross

We will quickly see if we can incorporate these features into a new model. As we are now going into more detail with the inter-relationship between given name and surname, it is probably no longer sensible to model them as independent comparisons, and so we will need to switch to a combined comparison on full name.

# we need to append a full name column to our source data frames
# so that we can use it for term frequency adjustments
dfs = [
    duckdb.sql("select *, given_name || '_' || surname as full_name from dataset")
    .arrow()
    .read_all()
    for dataset in dfs
]


extended_model_settings = {
    "unique_id_column_name": "rec_id",
    "link_type": "link_only",
    "blocking_rules_to_generate_predictions": blocking_rules,
    "comparisons": [
        {
            "output_column_name": "Full name",
            "comparison_levels": [
                {
                    "sql_condition": "(given_name_l IS NULL OR given_name_r IS NULL) and (surname_l IS NULL OR surname_r IS NULL)",
                    "label_for_charts": "Null",
                    "is_null_level": True,
                },
                # full name match
                cll.ExactMatchLevel("full_name", term_frequency_adjustments=True),
                # typos - keep levels across full name rather than scoring separately
                cll.JaroWinklerLevel("full_name", 0.9),
                cll.JaroWinklerLevel("full_name", 0.7),
                # name switched
                cll.ColumnsReversedLevel("given_name", "surname"),
                # name switched + typo
                {
                    "sql_condition": "jaro_winkler_similarity(given_name_l, surname_r) + jaro_winkler_similarity(surname_l, given_name_r) >= 1.8",
                    "label_for_charts": "switched + jaro_winkler_similarity >= 1.8",
                },
                {
                    "sql_condition": "jaro_winkler_similarity(given_name_l, surname_r) + jaro_winkler_similarity(surname_l, given_name_r) >= 1.4",
                    "label_for_charts": "switched + jaro_winkler_similarity >= 1.4",
                },
                # single name match
                cll.ExactMatchLevel("given_name", term_frequency_adjustments=True),
                cll.ExactMatchLevel("surname", term_frequency_adjustments=True),
                # single name cross-match
                {
                    "sql_condition": "given_name_l = surname_r OR surname_l = given_name_r",
                    "label_for_charts": "single name cross-matches",
                },  # single name typos
                cll.JaroWinklerLevel("given_name", 0.9),
                cll.JaroWinklerLevel("surname", 0.9),
                # the rest
                cll.ElseLevel(),
            ],
        },
        cl.DateOfBirthComparison(
            "date_of_birth",
            input_is_string=True,
            datetime_format="%Y%m%d",
            invalid_dates_as_null=True,
        ),
        {
            "output_column_name": "Social security ID",
            "comparison_levels": [
                cll.NullLevel("soc_sec_id"),
                cll.ExactMatchLevel("soc_sec_id", term_frequency_adjustments=True),
                cll.DamerauLevenshteinLevel("soc_sec_id", 1),
                cll.DamerauLevenshteinLevel("soc_sec_id", 2),
                cll.ElseLevel(),
            ],
        },
        {
            "output_column_name": "Street number",
            "comparison_levels": [
                cll.NullLevel("street_number"),
                cll.ExactMatchLevel("street_number", term_frequency_adjustments=True),
                cll.DamerauLevenshteinLevel("street_number", 1),
                cll.ElseLevel(),
            ],
        },
        {
            "output_column_name": "Postcode",
            "comparison_levels": [
                cll.NullLevel("postcode"),
                cll.ExactMatchLevel("postcode", term_frequency_adjustments=True),
                cll.DamerauLevenshteinLevel("postcode", 1),
                cll.DamerauLevenshteinLevel("postcode", 2),
                cll.ElseLevel(),
            ],
        },
        # we don't consider further location columns as they will be strongly correlated with postcode
    ],
    "retain_intermediate_calculation_columns": True,
}
# train
db_api = DuckDBAPI()
dfs_sdf = [db_api.register(df) for df in dfs]
linker_advanced = Linker(dfs_sdf, extended_model_settings)
linker_advanced.training.estimate_probability_two_random_records_match(
    deterministic_rules, recall=0.8
)
# We recommend increasing target rows to 1e8 improve accuracy for u
# values in full name comparison, as we have subdivided the data more finely

# Here, 1e7 for speed
linker_advanced.training.estimate_u_using_random_sampling(max_pairs=1e7)
Probability two random records match is estimated to be  0.000239.
This means that amongst all possible pairwise record comparisons, one in 4,185.85 are expected to match.  With 25,000,000 total possible comparisons, we expect a total of around 5,972.50 matching pairs


----- Estimating u probabilities using random sampling -----


Estimating u with: max_pairs = 10,000,000, min_count_per_level = 100, num_chunks = 10



Estimating u for: Full name (Comparison 1 of 5)


  Running probe chunk (~1.00% of max_pairs)


  Min u_count: 0 for comparison level single name cross-matches (cvv=3)


  Probe did not converge; restarting with normal chunking



  Running chunk 1/10


  Count of 0 for level single name cross-matches (cvv=3). Chunk took 1.5 seconds.


  Min u_count not hit, continuing.


  Running chunk 2/10


  Count of 0 for level single name cross-matches (cvv=3). Chunk took 1.2 seconds.


  Min u_count not hit, continuing.


  Running chunk 3/10


  Count of 0 for level single name cross-matches (cvv=3). Chunk took 1.2 seconds.


  Min u_count not hit, continuing.


  Running chunk 4/10


  Count of 0 for level single name cross-matches (cvv=3). Chunk took 1.1 seconds.


  Min u_count not hit, continuing.


  Running chunk 5/10


  Count of 0 for level single name cross-matches (cvv=3). Chunk took 1.2 seconds.


  Min u_count not hit, continuing.


  Running chunk 6/10


  Count of 0 for level single name cross-matches (cvv=3). Chunk took 1.2 seconds.


  Min u_count not hit, continuing.


  Running chunk 7/10


  Count of 0 for level single name cross-matches (cvv=3). Chunk took 1.3 seconds.


  Min u_count not hit, continuing.


  Running chunk 8/10


  Count of 1 for level single name cross-matches (cvv=3). Chunk took 1.2 seconds.


  Min u_count not hit, continuing.


  Running chunk 9/10


  Count of 1 for level single name cross-matches (cvv=3). Chunk took 1.2 seconds.


  Min u_count not hit, continuing.


  Running chunk 10/10


  Count of 1 for level single name cross-matches (cvv=3). Chunk took 1.2 seconds.


  Min u_count not hit, continuing.



Estimating u for: date_of_birth (Comparison 2 of 5)


  Running probe chunk (~1.00% of max_pairs)


  Min u_count: 29 for comparison level Exact match on date of birth (cvv=5)


  Probe did not converge; restarting with normal chunking



  Running chunk 1/10


  Count of 203 for level Exact match on date of birth (cvv=5). Chunk took 0.8 seconds.


  Exiting early since min count of 203 exceeds min_count_per_level = 100



Estimating u for: Social security ID (Comparison 3 of 5)


  Running probe chunk (~1.00% of max_pairs)


  Min u_count: 6 for comparison level Damerau-Levenshtein distance of soc_sec_id <= 1 (cvv=2)


  Probe did not converge; restarting with normal chunking



  Running chunk 1/10


  Count of 19 for level Damerau-Levenshtein distance of soc_sec_id <= 1 (cvv=2). Chunk took 1.7 seconds.


  Min u_count not hit, continuing.


  Running chunk 2/10


  Count of 40 for level Damerau-Levenshtein distance of soc_sec_id <= 1 (cvv=2). Chunk took 1.7 seconds.


  Min u_count not hit, continuing.


  Running chunk 3/10


  Count of 59 for level Damerau-Levenshtein distance of soc_sec_id <= 1 (cvv=2). Chunk took 1.8 seconds.


  Min u_count not hit, continuing.


  Running chunk 4/10


  Count of 78 for level Damerau-Levenshtein distance of soc_sec_id <= 1 (cvv=2). Chunk took 1.7 seconds.


  Min u_count not hit, continuing.


  Running chunk 5/10


  Count of 95 for level Damerau-Levenshtein distance of soc_sec_id <= 1 (cvv=2). Chunk took 1.8 seconds.


  Min u_count not hit, continuing.


  Running chunk 6/10


  Count of 109 for level Damerau-Levenshtein distance of soc_sec_id <= 1 (cvv=2). Chunk took 1.7 seconds.


  Exiting early since min count of 109 exceeds min_count_per_level = 100



Estimating u for: Street number (Comparison 4 of 5)


  Running probe chunk (~1.00% of max_pairs)


  Min u_count: 2,020 for comparison level Exact match on street_number (cvv=2)


  Exiting early since min count of 2,020 exceeds min_count_per_level = 100



Estimating u for: Postcode (Comparison 5 of 5)


  Running probe chunk (~1.00% of max_pairs)


  Min u_count: 156 for comparison level Exact match on postcode (cvv=3)


  Exiting early since min count of 156 exceeds min_count_per_level = 100



Estimated u probabilities using random sampling



Your model is not yet fully trained. Missing estimates for:
    - Full name (no m values are trained).
    - date_of_birth (no m values are trained).
    - Social security ID (no m values are trained).
    - Street number (no m values are trained).
    - Postcode (no m values are trained).
session_dob = (
    linker_advanced.training.estimate_parameters_using_expectation_maximisation(
        "l.date_of_birth = r.date_of_birth", estimate_without_term_frequencies=True
    )
)
----- Starting EM training session -----



[EM sampling] max_pairs is None — no sampling will be applied


Estimating the m probabilities of the model by blocking on:
l.date_of_birth = r.date_of_birth

Parameter estimates will be made for the following comparison(s):
    - Full name
    - Social security ID
    - Street number
    - Postcode

Parameter estimates cannot be made for the following comparison(s) since they are used in the blocking rules: 
    - date_of_birth





WARNING:
Level single name cross-matches on comparison Full name not observed in dataset, unable to train m value



Iteration 1: Largest change in params was -0.465 in the m_probability of Full name, level `Exact match on full_name`


Iteration 2: Largest change in params was 0.00251 in the m_probability of Social security ID, level `All other comparisons`


Iteration 3: Largest change in params was 4.85e-05 in the m_probability of Social security ID, level `All other comparisons`



EM converged after 3 iterations


m probability not trained for Full name - single name cross-matches (comparison vector value: 3). This usually means the comparison level was never observed in the training data.



Your model is not yet fully trained. Missing estimates for:
    - Full name (some m values are not trained).
    - date_of_birth (no m values are trained).
session_pc = (
    linker_advanced.training.estimate_parameters_using_expectation_maximisation(
        "l.postcode = r.postcode", estimate_without_term_frequencies=True
    )
)
----- Starting EM training session -----



[EM sampling] max_pairs is None — no sampling will be applied


Estimating the m probabilities of the model by blocking on:
l.postcode = r.postcode

Parameter estimates will be made for the following comparison(s):
    - Full name
    - date_of_birth
    - Social security ID
    - Street number

Parameter estimates cannot be made for the following comparison(s) since they are used in the blocking rules: 
    - Postcode





WARNING:
Level single name cross-matches on comparison Full name not observed in dataset, unable to train m value



Iteration 1: Largest change in params was 0.0284 in the m_probability of date_of_birth, level `All other comparisons`


Iteration 2: Largest change in params was 0.000418 in the m_probability of date_of_birth, level `All other comparisons`


Iteration 3: Largest change in params was 1.14e-05 in the m_probability of Social security ID, level `All other comparisons`



EM converged after 3 iterations


m probability not trained for Full name - single name cross-matches (comparison vector value: 3). This usually means the comparison level was never observed in the training data.



Your model is not yet fully trained. Missing estimates for:
    - Full name (some m values are not trained).
linker_advanced.visualisations.parameter_estimate_comparisons_chart()
linker_advanced.visualisations.match_weights_chart()
predictions_adv = linker_advanced.inference.predict()
clusters_adv = linker_advanced.clustering.cluster_pairwise_predictions_at_threshold(
    predictions_adv, threshold_match_probability=0.99
)
cluster_size_distribution_adv = clusters_adv.query_sql(
    """
    select cluster_size, count(*) as cluster_count
    from (
        select cluster_id, count(*) as cluster_size
        from {this}
        group by cluster_id
    )
    group by cluster_size
    order by cluster_size
    """
)
cluster_size_distribution_adv.as_duckdbpyrelation().show()
Blocking time: 0.04 seconds


Predict time (post-blocking): 0.50 seconds



 -- WARNING --
You have called predict(), but there are some parameter estimates which have neither been estimated or specified in your settings dictionary.  To produce predictions the following untrained parameters will use default values.
Comparison: 'Full name':
    m values not fully trained


Completed iteration 1, num edges remaining to process: 0


┌──────────────┬───────────────┐
│ cluster_size │ cluster_count │
│    int64     │     int64     │
├──────────────┼───────────────┤
│            1 │            80 │
│            2 │          4960 │
└──────────────┴───────────────┘

This is a pretty modest improvement on our previous model - however it is worth re-iterating that we should not necessarily expect to recover all matches, as in several cases it may be unreasonable for a model to have reasonable confidence that two records refer to the same entity.

If we wished to improve matters we could iterate on this process - investigating where our model is not performing as we would hope, and seeing how we can adjust these areas to address these shortcomings.