Skip to content

Methods in Linker.inference

Use your Splink model to make predictions (perform inference). Accessed via linker.inference.

compute_blocked_pairs_for_predict()

Compute blocked pairs for prediction.

Uses the blocking rules specified in the blocking_rules_to_generate_predictions key of the settings to generate the candidate pairs that predict() would score.

This is useful when you want to materialise blocked pairs separately from scoring, for example to write them out and re-register them in a different job or on a different machine:

blocked_pairs = linker.inference.compute_blocked_pairs_for_predict()
# Write the blocked pairs out to parquet (DuckDB example)
blocked_pairs.as_duckdbpyrelation().to_parquet("blocked_pairs.parquet")

# Then in a different session e.g. on a different machine
blocked_pairs = db_api.register("blocked_pairs.parquet")
linker.table_management.register_blocked_pairs_for_predict(blocked_pairs)
predictions = linker.inference.predict()

To compute blocked pairs for a single chunk instead, use compute_blocked_pairs_for_predict_chunk().

Returns:

Name Type Description
SplinkDataFrame SplinkDataFrame

The blocked pairs table, also stored in cache.

Examples:

blocked_pairs = linker.inference.compute_blocked_pairs_for_predict()

compute_blocked_pairs_for_predict_chunk(left_chunk=None, right_chunk=None)

Compute blocked pairs for a single chunk of the prediction.

Uses the blocking rules specified in the blocking_rules_to_generate_predictions key of the settings to generate the candidate pairs for one slice of the data that predict_chunk() would score.

This is useful when you want to materialise blocked pairs for a single chunk separately from scoring, for example to distribute the work across many machines:

blocked_pairs = linker.inference.compute_blocked_pairs_for_predict_chunk(
    left_chunk=(1, 3),
    right_chunk=(2, 4),
)
# Write the blocked pairs out to parquet (DuckDB example)
blocked_pairs.as_duckdbpyrelation().to_parquet("blocked_pairs.parquet")

# Then in a different session e.g. on a different machine
blocked_pairs = db_api.register("blocked_pairs.parquet")
linker.table_management.register_blocked_pairs_for_predict(blocked_pairs)
predictions = linker.inference.predict()

To compute the full blocked-pairs table in one go instead, use compute_blocked_pairs_for_predict().

Parameters:

Name Type Description Default
left_chunk tuple[int, int] | None

Optional tuple of (chunk_number, total_chunks) for filtering left side records. For example, (1, 3) means chunk 1 of 3.

None
right_chunk tuple[int, int] | None

Optional tuple of (chunk_number, total_chunks) for filtering right side records. For example, (2, 4) means chunk 2 of 4.

None

Returns:

Name Type Description
SplinkDataFrame SplinkDataFrame

The blocked pairs table, also stored in cache.

Examples:

linker.inference.compute_blocked_pairs_for_predict_chunk(
    left_chunk=(1, 3),
    right_chunk=(2, 4),
)

linker.inference.compute_blocked_pairs_for_predict_chunk(
    left_chunk=(1, 1),
    right_chunk=(1, 1),
)

Uses the blocking rules specified by blocking_rules_to_generate_predictions in your settings to generate pairwise record comparisons.

For deterministic linkage, this should be a list of blocking rules which are strict enough to generate only true links.

Deterministic linkage, however, is likely to result in missed links (false negatives).

Returns:

Name Type Description
SplinkDataFrame SplinkDataFrame

A SplinkDataFrame of the pairwise comparisons.

```py
settings = SettingsCreator(
    link_type="dedupe_only",
    blocking_rules_to_generate_predictions=[
        block_on("first_name", "surname"),
        block_on("dob", "first_name"),
    ],
)

df = db_api.register(df, dataset_display_name="input_table")
linker = Linker(df, settings)
splink_df = linker.inference.deterministic_link()
```

predict(threshold_match_probability=None, threshold_match_weight=None, num_chunks_left=None, num_chunks_right=None, warning_mode='auto')

Create a dataframe of scored pairwise comparisons using the parameters of the linkage model.

Uses the blocking rules specified in the blocking_rules_to_generate_predictions key of the settings to generate the pairwise comparisons.

If blocked pairs have been manually registered using linker.table_management.register_blocked_pairs_for_predict(), this method scores exactly that registered table. In that workflow the chunking arguments (num_chunks_left / num_chunks_right) are not supported, because Splink cannot own chunking of a table you have already materialised.

Parameters:

Name Type Description Default
threshold_match_probability float

If specified, filter the results to include only pairwise comparisons with a match_probability above this threshold. Defaults to None.

None
threshold_match_weight float

If specified, filter the results to include only pairwise comparisons with a match_weight above this threshold. Defaults to None.

None
num_chunks_left int

If specified along with num_chunks_right, the prediction will be split into chunks and processed iteratively. This can help manage memory usage for large datasets.

None
num_chunks_right int

If specified along with num_chunks_left, the prediction will be split into chunks and processed iteratively.

None
warning_mode str

Control emission of the warning shown when predict runs with untrained model parameters. Use "auto" to emit once per call to predict(), "always" to force emission once per call to predict(), or "never" to suppress this warning.

'auto'

Examples:

df = db_api.register(df, dataset_display_name="input_table")
linker = Linker(df, "saved_settings.json")
splink_df = linker.inference.predict(threshold_match_probability=0.95)
splink_df.as_pandas_dataframe(limit=5)

# With chunking for large datasets
splink_df = linker.inference.predict(
    threshold_match_probability=0.95,
    num_chunks_left=3,
    num_chunks_right=4
)

predict_chunk(left_chunk=None, right_chunk=None, threshold_match_probability=None, threshold_match_weight=None, warning_mode='auto')

Create a dataframe of scored pairwise comparisons for a specific chunk of the data.

This method lets Splink compute and score blocking for a single slice of the data, for example one worker per slice in a distributed run. It is not supported when blocked pairs have been manually registered using linker.table_management.register_blocked_pairs_for_predict(); in that workflow call linker.inference.predict() to score the registered table.

Parameters:

Name Type Description Default
left_chunk tuple[int, int]

Tuple of (chunk_number, total_chunks) for filtering left side records. For example, (1, 3) means chunk 1 of 3.

None
right_chunk tuple[int, int]

Tuple of (chunk_number, total_chunks) for filtering right side records. For example, (2, 4) means chunk 2 of 4.

None
threshold_match_probability float

If specified, filter the results to include only pairwise comparisons with a match_probability above this threshold. Defaults to None.

None
threshold_match_weight float

If specified, filter the results to include only pairwise comparisons with a match_weight above this threshold. Defaults to None.

None
warning_mode str

Control emission of the warning shown when predict runs with untrained model parameters. Use "auto" to emit once per direct call to predict_chunk(), "always" to force emission, or "never" to suppress this warning.

'auto'

Examples:

df = db_api.register(df, dataset_display_name="input_table")
linker = Linker(df, "saved_settings.json")
# Process chunk 1 of 3 on left, chunk 2 of 4 on right
splink_df = linker.inference.predict_chunk(
    left_chunk=(1, 3),
    right_chunk=(2, 4),
    threshold_match_probability=0.5
)
splink_df.as_pandas_dataframe(limit=5)

splink_df = linker.inference.predict_chunk(
    left_chunk=(1, 1),
    right_chunk=(1, 1),
)

score_pair(record_left, record_right, include_found_by_blocking_rules=False)

Use the linkage model to score a single pairwise record comparison.

Each input may be a Python dict representing a single record, or a SplinkDataFrame.

The usual usage is to provide any required term frequency values directly in the input records as hardcoded term frequency columns (e.g. a tf_first_name column). If these values are not provided, Splink falls back to any registered term frequency lookup tables, or term frequency values derived from the input data.

Parameters:

Name Type Description Default
record_left dict | SplinkDataFrame

the left-hand record. Column names and data types must be the same as the columns in the settings object

required
record_right dict | SplinkDataFrame

the right-hand record. Column names and data types must be the same as the columns in the settings object

required
include_found_by_blocking_rules bool

If True, outputs a column indicating whether the record pair would have been found by any of the blocking rules specified in settings.blocking_rules_to_generate_predictions. Defaults to False.

False

Examples:

df = db_api.register(df, dataset_display_name="input_table")
linker = Linker(df, "saved_settings.json")

# If you do not provide tf values in the records, you should load or
# pre-compute tf tables for any columns with term frequency adjustments
linker.table_management.compute_tf_table("first_name")
# OR
linker.table_management.register_term_frequency_lookup(tf, "first_name")

record_1 = {'unique_id': 1,
    'first_name': "John",
    'surname': "Smith",
    'dob': "1971-05-24",
    'city': "London",
    'email': "john@smith.net",
    'tf_first_name': 0.001,
    }

record_2 = {'unique_id': 1,
    'first_name': "Jon",
    'surname': "Smith",
    'dob': "1971-05-23",
    'city': "London",
    'email': "john@smith.net",
    'tf_first_name': 0.0005,
    }
df = linker.inference.score_pair(record_1, record_2)

Returns:

Name Type Description
SplinkDataFrame SplinkDataFrame

Pairwise comparison with scored prediction

score_pairs(records_left, records_right, include_found_by_blocking_rules=False)

Use the linkage model to score pairwise record comparisons formed from the cartesian product of the two inputs provided.

No blocking rules are applied: every record in records_left is compared against every record in records_right. To score a single known pair use score_pair; to generate candidate pairs using blocking rules, use predict_within or predict_between instead.

Each input may be a list of Python dicts representing records, or a SplinkDataFrame.

The usual usage is to provide any required term frequency values directly in the input records as hardcoded term frequency columns (e.g. a tf_first_name column). If these values are not provided, Splink falls back to any registered term frequency lookup tables, or term frequency values derived from the input data.

Parameters:

Name Type Description Default
records_left list[dict] | SplinkDataFrame

the left-hand records. Column names and data types must be the same as the columns in the settings object

required
records_right list[dict] | SplinkDataFrame

the right-hand records. Column names and data types must be the same as the columns in the settings object

required
include_found_by_blocking_rules bool

If True, outputs a column indicating whether the record pair would have been found by any of the blocking rules specified in settings.blocking_rules_to_generate_predictions. Defaults to False.

False

Examples:

df = db_api.register(df, dataset_display_name="input_table")
linker = Linker(df, "saved_settings.json")

linker.table_management.compute_tf_table("first_name")

df = linker.inference.score_pairs(records_left, records_right)

Returns:

Name Type Description
SplinkDataFrame SplinkDataFrame

Pairwise comparisons with scored predictions

predict_within(splink_dataframe_or_dataframes, *, link_type=None, blocking_rules_to_generate_predictions=None, threshold_match_probability=None, threshold_match_weight=None, warning_mode='auto')

Generate blocked, scored pairwise predictions within a supplied collection of records, using the trained model.

The input shape mirrors the Linker constructor: pass a single SplinkDataFrame for dedupe_only, or a list of SplinkDataFrames for link_only and link_and_dedupe. Candidate pairs are generated using blocking rules (the trained rules by default), respecting the model's link_type, and scored.

Unlike predict() this does not derive term-frequency values from the supplied data: any required term-frequency tables must be registered (or supplied as hardcoded tf_* columns), otherwise a SplinkException is raised.

Parameters:

Name Type Description Default
splink_dataframe_or_dataframes SplinkDataFrame | Sequence[SplinkDataFrame]

A single SplinkDataFrame or a list of them, registered against the same db_api.

required
link_type LinkTypeLiteralType | None

Optionally override the trained link_type.

None
blocking_rules_to_generate_predictions list[BlockingRuleCreator | BlockingRule | str | dict[str, Any]] | None

Optionally override the blocking rules used to generate candidate pairs.

None
threshold_match_probability float | None

If specified, only return pairs with a match probability above this threshold.

None
threshold_match_weight float | None

If specified, only return pairs with a match weight above this threshold.

None
warning_mode PredictUntrainedWarningMode

Control emission of the untrained-model warning.

'auto'

Returns:

Name Type Description
SplinkDataFrame SplinkDataFrame

A SplinkDataFrame of the scored pairwise comparisons.

predict_between(left, right, *, link_type=None, blocking_rules_to_generate_predictions=None, threshold_match_probability=None, threshold_match_weight=None, warning_mode='auto')

Generate blocked, scored pairwise predictions between two supplied collections of records, using the trained model.

Candidate pairs are generated between records in the left dataset(s) and right dataset(s) respectively, never within them. The trained model's link_type source condition is then also applied, e.g. for link_only, pairs are additionally required to come from different source datasets

One key use case is incremental record linkage, in which we want to find links between the existing and new records, but not within the existing records.

Often, in addition, you'd want to find links within the new records, for which you'd need to also use predict_within().

Note that left and right are roles (for example existing vs new), not source datasets

Term-frequency tables must be registered (or supplied as hardcoded tf_* columns), otherwise a SplinkException is raised.

Parameters:

Name Type Description Default
left SplinkDataFrame | Sequence[SplinkDataFrame]

The left-hand (role) collection: a single SplinkDataFrame or a sequence of them.

required
right SplinkDataFrame | Sequence[SplinkDataFrame]

The right-hand (role) collection: a single SplinkDataFrame or a sequence of them.

required
link_type LinkTypeLiteralType | None

Optionally override the trained link_type.

None
blocking_rules_to_generate_predictions list[BlockingRuleCreator | BlockingRule | str | dict[str, Any]] | None

Optionally override the blocking rules used to generate candidate pairs.

None
threshold_match_probability float | None

If specified, only return pairs with a match probability above this threshold.

None
threshold_match_weight float | None

If specified, only return pairs with a match weight above this threshold.

None
warning_mode PredictUntrainedWarningMode

Control emission of the untrained-model warning.

'auto'

Returns:

Name Type Description
SplinkDataFrame SplinkDataFrame

A SplinkDataFrame of the scored pairwise comparisons.