Methods in Linker.inference¶
Use your Splink model to make predictions (perform inference). Accessed via
linker.inference.
compute_blocked_pairs_for_predict()
¶
Compute blocked pairs for prediction.
Uses the blocking rules specified in the
blocking_rules_to_generate_predictions key of the settings to generate
the candidate pairs that predict() would score.
This is useful when you want to materialise blocked pairs separately from scoring, for example to write them out and re-register them in a different job or on a different machine:
blocked_pairs = linker.inference.compute_blocked_pairs_for_predict()
# Write the blocked pairs out to parquet (DuckDB example)
blocked_pairs.as_duckdbpyrelation().to_parquet("blocked_pairs.parquet")
# Then in a different session e.g. on a different machine
blocked_pairs = db_api.register("blocked_pairs.parquet")
linker.table_management.register_blocked_pairs_for_predict(blocked_pairs)
predictions = linker.inference.predict()
To compute blocked pairs for a single chunk instead, use
compute_blocked_pairs_for_predict_chunk().
Returns:
| Name | Type | Description |
|---|---|---|
SplinkDataFrame |
SplinkDataFrame
|
The blocked pairs table, also stored in cache. |
Examples:
blocked_pairs = linker.inference.compute_blocked_pairs_for_predict()
compute_blocked_pairs_for_predict_chunk(left_chunk=None, right_chunk=None)
¶
Compute blocked pairs for a single chunk of the prediction.
Uses the blocking rules specified in the
blocking_rules_to_generate_predictions key of the settings to generate
the candidate pairs for one slice of the data that predict_chunk() would
score.
This is useful when you want to materialise blocked pairs for a single chunk separately from scoring, for example to distribute the work across many machines:
blocked_pairs = linker.inference.compute_blocked_pairs_for_predict_chunk(
left_chunk=(1, 3),
right_chunk=(2, 4),
)
# Write the blocked pairs out to parquet (DuckDB example)
blocked_pairs.as_duckdbpyrelation().to_parquet("blocked_pairs.parquet")
# Then in a different session e.g. on a different machine
blocked_pairs = db_api.register("blocked_pairs.parquet")
linker.table_management.register_blocked_pairs_for_predict(blocked_pairs)
predictions = linker.inference.predict()
To compute the full blocked-pairs table in one go instead, use
compute_blocked_pairs_for_predict().
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
left_chunk
|
tuple[int, int] | None
|
Optional tuple of (chunk_number, total_chunks) for filtering left side records. For example, (1, 3) means chunk 1 of 3. |
None
|
right_chunk
|
tuple[int, int] | None
|
Optional tuple of (chunk_number, total_chunks) for filtering right side records. For example, (2, 4) means chunk 2 of 4. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
SplinkDataFrame |
SplinkDataFrame
|
The blocked pairs table, also stored in cache. |
Examples:
linker.inference.compute_blocked_pairs_for_predict_chunk(
left_chunk=(1, 3),
right_chunk=(2, 4),
)
linker.inference.compute_blocked_pairs_for_predict_chunk(
left_chunk=(1, 1),
right_chunk=(1, 1),
)
deterministic_link()
¶
Uses the blocking rules specified by
blocking_rules_to_generate_predictions in your settings to
generate pairwise record comparisons.
For deterministic linkage, this should be a list of blocking rules which are strict enough to generate only true links.
Deterministic linkage, however, is likely to result in missed links (false negatives).
Returns:
| Name | Type | Description |
|---|---|---|
SplinkDataFrame |
SplinkDataFrame
|
A SplinkDataFrame of the pairwise comparisons. |
```py
settings = SettingsCreator(
link_type="dedupe_only",
blocking_rules_to_generate_predictions=[
block_on("first_name", "surname"),
block_on("dob", "first_name"),
],
)
df = db_api.register(df, dataset_display_name="input_table")
linker = Linker(df, settings)
splink_df = linker.inference.deterministic_link()
```
predict(threshold_match_probability=None, threshold_match_weight=None, num_chunks_left=None, num_chunks_right=None, warning_mode='auto')
¶
Create a dataframe of scored pairwise comparisons using the parameters of the linkage model.
Uses the blocking rules specified in the
blocking_rules_to_generate_predictions key of the settings to
generate the pairwise comparisons.
If blocked pairs have been manually registered using
linker.table_management.register_blocked_pairs_for_predict(), this
method scores exactly that registered table. In that workflow the
chunking arguments (num_chunks_left / num_chunks_right) are not
supported, because Splink cannot own chunking of a table you have
already materialised.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
threshold_match_probability
|
float
|
If specified, filter the results to include only pairwise comparisons with a match_probability above this threshold. Defaults to None. |
None
|
threshold_match_weight
|
float
|
If specified, filter the results to include only pairwise comparisons with a match_weight above this threshold. Defaults to None. |
None
|
num_chunks_left
|
int
|
If specified along with num_chunks_right, the prediction will be split into chunks and processed iteratively. This can help manage memory usage for large datasets. |
None
|
num_chunks_right
|
int
|
If specified along with num_chunks_left, the prediction will be split into chunks and processed iteratively. |
None
|
warning_mode
|
str
|
Control emission of the warning shown when
predict runs with untrained model parameters. Use "auto" to emit
once per call to |
'auto'
|
Examples:
df = db_api.register(df, dataset_display_name="input_table")
linker = Linker(df, "saved_settings.json")
splink_df = linker.inference.predict(threshold_match_probability=0.95)
splink_df.as_pandas_dataframe(limit=5)
# With chunking for large datasets
splink_df = linker.inference.predict(
threshold_match_probability=0.95,
num_chunks_left=3,
num_chunks_right=4
)
predict_chunk(left_chunk=None, right_chunk=None, threshold_match_probability=None, threshold_match_weight=None, warning_mode='auto')
¶
Create a dataframe of scored pairwise comparisons for a specific chunk of the data.
This method lets Splink compute and score blocking for a single slice of
the data, for example one worker per slice in a distributed run. It is not
supported when blocked pairs have been manually registered using
linker.table_management.register_blocked_pairs_for_predict(); in that
workflow call linker.inference.predict() to score the registered table.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
left_chunk
|
tuple[int, int]
|
Tuple of (chunk_number, total_chunks) for filtering left side records. For example, (1, 3) means chunk 1 of 3. |
None
|
right_chunk
|
tuple[int, int]
|
Tuple of (chunk_number, total_chunks) for filtering right side records. For example, (2, 4) means chunk 2 of 4. |
None
|
threshold_match_probability
|
float
|
If specified, filter the results to include only pairwise comparisons with a match_probability above this threshold. Defaults to None. |
None
|
threshold_match_weight
|
float
|
If specified, filter the results to include only pairwise comparisons with a match_weight above this threshold. Defaults to None. |
None
|
warning_mode
|
str
|
Control emission of the warning shown when
predict runs with untrained model parameters. Use "auto" to emit
once per direct call to |
'auto'
|
Examples:
df = db_api.register(df, dataset_display_name="input_table")
linker = Linker(df, "saved_settings.json")
# Process chunk 1 of 3 on left, chunk 2 of 4 on right
splink_df = linker.inference.predict_chunk(
left_chunk=(1, 3),
right_chunk=(2, 4),
threshold_match_probability=0.5
)
splink_df.as_pandas_dataframe(limit=5)
splink_df = linker.inference.predict_chunk(
left_chunk=(1, 1),
right_chunk=(1, 1),
)
score_pair(record_left, record_right, include_found_by_blocking_rules=False)
¶
Use the linkage model to score a single pairwise record comparison.
Each input may be a Python dict representing a single record, or a
SplinkDataFrame.
The usual usage is to provide any required term frequency values directly in the input records as hardcoded term frequency columns (e.g. a tf_first_name column). If these values are not provided, Splink falls back to any registered term frequency lookup tables, or term frequency values derived from the input data.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
record_left
|
dict | SplinkDataFrame
|
the left-hand record. Column names and data types must be the same as the columns in the settings object |
required |
record_right
|
dict | SplinkDataFrame
|
the right-hand record. Column names and data types must be the same as the columns in the settings object |
required |
include_found_by_blocking_rules
|
bool
|
If True, outputs a column indicating whether the record pair would have been found by any of the blocking rules specified in settings.blocking_rules_to_generate_predictions. Defaults to False. |
False
|
Examples:
df = db_api.register(df, dataset_display_name="input_table")
linker = Linker(df, "saved_settings.json")
# If you do not provide tf values in the records, you should load or
# pre-compute tf tables for any columns with term frequency adjustments
linker.table_management.compute_tf_table("first_name")
# OR
linker.table_management.register_term_frequency_lookup(tf, "first_name")
record_1 = {'unique_id': 1,
'first_name': "John",
'surname': "Smith",
'dob': "1971-05-24",
'city': "London",
'email': "john@smith.net",
'tf_first_name': 0.001,
}
record_2 = {'unique_id': 1,
'first_name': "Jon",
'surname': "Smith",
'dob': "1971-05-23",
'city': "London",
'email': "john@smith.net",
'tf_first_name': 0.0005,
}
df = linker.inference.score_pair(record_1, record_2)
Returns:
| Name | Type | Description |
|---|---|---|
SplinkDataFrame |
SplinkDataFrame
|
Pairwise comparison with scored prediction |
score_pairs(records_left, records_right, include_found_by_blocking_rules=False)
¶
Use the linkage model to score pairwise record comparisons formed from the cartesian product of the two inputs provided.
No blocking rules are applied: every record in records_left is compared
against every record in records_right. To score a single known pair use
score_pair; to generate candidate pairs using blocking rules, use
predict_within or predict_between instead.
Each input may be a list of Python dicts representing records, or a
SplinkDataFrame.
The usual usage is to provide any required term frequency values directly in the input records as hardcoded term frequency columns (e.g. a tf_first_name column). If these values are not provided, Splink falls back to any registered term frequency lookup tables, or term frequency values derived from the input data.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
records_left
|
list[dict] | SplinkDataFrame
|
the left-hand records. Column names and data types must be the same as the columns in the settings object |
required |
records_right
|
list[dict] | SplinkDataFrame
|
the right-hand records. Column names and data types must be the same as the columns in the settings object |
required |
include_found_by_blocking_rules
|
bool
|
If True, outputs a column indicating whether the record pair would have been found by any of the blocking rules specified in settings.blocking_rules_to_generate_predictions. Defaults to False. |
False
|
Examples:
df = db_api.register(df, dataset_display_name="input_table")
linker = Linker(df, "saved_settings.json")
linker.table_management.compute_tf_table("first_name")
df = linker.inference.score_pairs(records_left, records_right)
Returns:
| Name | Type | Description |
|---|---|---|
SplinkDataFrame |
SplinkDataFrame
|
Pairwise comparisons with scored predictions |
predict_within(splink_dataframe_or_dataframes, *, link_type=None, blocking_rules_to_generate_predictions=None, threshold_match_probability=None, threshold_match_weight=None, warning_mode='auto')
¶
Generate blocked, scored pairwise predictions within a supplied collection of records, using the trained model.
The input shape mirrors the Linker constructor: pass a single
SplinkDataFrame for dedupe_only, or a list of
SplinkDataFrames for link_only and link_and_dedupe. Candidate pairs
are generated using blocking rules (the trained rules by default), respecting
the model's link_type, and scored.
Unlike predict() this does not derive term-frequency values from the
supplied data: any required term-frequency tables must be registered (or
supplied as hardcoded tf_* columns), otherwise a SplinkException is
raised.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
splink_dataframe_or_dataframes
|
SplinkDataFrame | Sequence[SplinkDataFrame]
|
A single |
required |
link_type
|
LinkTypeLiteralType | None
|
Optionally override the trained |
None
|
blocking_rules_to_generate_predictions
|
list[BlockingRuleCreator | BlockingRule | str | dict[str, Any]] | None
|
Optionally override the blocking rules used to generate candidate pairs. |
None
|
threshold_match_probability
|
float | None
|
If specified, only return pairs with a match probability above this threshold. |
None
|
threshold_match_weight
|
float | None
|
If specified, only return pairs with a match weight above this threshold. |
None
|
warning_mode
|
PredictUntrainedWarningMode
|
Control emission of the untrained-model warning. |
'auto'
|
Returns:
| Name | Type | Description |
|---|---|---|
SplinkDataFrame |
SplinkDataFrame
|
A SplinkDataFrame of the scored pairwise comparisons. |
predict_between(left, right, *, link_type=None, blocking_rules_to_generate_predictions=None, threshold_match_probability=None, threshold_match_weight=None, warning_mode='auto')
¶
Generate blocked, scored pairwise predictions between two supplied collections of records, using the trained model.
Candidate pairs are generated between records in the left dataset(s)
and right dataset(s) respectively, never within them.
The trained model's link_type source
condition is then also applied, e.g. for link_only, pairs are additionally
required to come from different source datasets
One key use case is incremental record linkage, in which we want to find links between the existing and new records, but not within the existing records.
Often, in addition, you'd want to find links within the new records, for which you'd need to also use predict_within().
Note that left and right are roles (for example existing vs new),
not source datasets
Term-frequency tables must be registered (or
supplied as hardcoded tf_* columns), otherwise a SplinkException is
raised.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
left
|
SplinkDataFrame | Sequence[SplinkDataFrame]
|
The left-hand (role) collection: a single |
required |
right
|
SplinkDataFrame | Sequence[SplinkDataFrame]
|
The right-hand (role) collection: a single |
required |
link_type
|
LinkTypeLiteralType | None
|
Optionally override the trained |
None
|
blocking_rules_to_generate_predictions
|
list[BlockingRuleCreator | BlockingRule | str | dict[str, Any]] | None
|
Optionally override the blocking rules used to generate candidate pairs. |
None
|
threshold_match_probability
|
float | None
|
If specified, only return pairs with a match probability above this threshold. |
None
|
threshold_match_weight
|
float | None
|
If specified, only return pairs with a match weight above this threshold. |
None
|
warning_mode
|
PredictUntrainedWarningMode
|
Control emission of the untrained-model warning. |
'auto'
|
Returns:
| Name | Type | Description |
|---|---|---|
SplinkDataFrame |
SplinkDataFrame
|
A SplinkDataFrame of the scored pairwise comparisons. |