Skip to content

Documentation forsplink.blocking_analysis

chart_comparisons_from_blocking_rules(splink_dataframe_or_dataframes, *, blocking_rules, link_type, unique_id_column_name='unique_id', source_dataset_column_name=None, record_sample_proportion=0.05)

Produce a chart of the cumulative number of comparisons generated by one or more blocking rules.

See count_comparisons_from_blocking_rules for details of the underlying computation and the meaning of record_sample_proportion.

Parameters:

Name Type Description Default
splink_dataframe_or_dataframes SplinkDataFrame | Sequence[SplinkDataFrame]

Input data

required
blocking_rules BlockingRuleLike | Iterable[BlockingRuleLike]

A single blocking rule or an iterable of blocking rules to analyse.

required
link_type user_input_link_type_options

The link type - "link_only", "dedupe_only" or "link_and_dedupe"

required
unique_id_column_name str

Defaults to "unique_id".

'unique_id'
source_dataset_column_name Optional[str]

Defaults to None.

None
record_sample_proportion float

Defaults to 0.05, meaning Splink estimates the count from a deterministic 5% sample of input records on each side of the blocking join. Set record_sample_proportion=1.0 to compute exact counts.

0.05

Returns:

Name Type Description
CumulativeBlockingRuleComparisonsGeneratedChart CumulativeBlockingRuleComparisonsGeneratedChart

The chart.

count_comparisons_from_blocking_rules(splink_dataframe_or_dataframes, *, blocking_rules, link_type, unique_id_column_name='unique_id', source_dataset_column_name=None, record_sample_proportion=0.05)

Analyse one or more blocking rules to understand how many comparisons they will generate.

The comparisons are counted by actually creating the blocked pairs (post filter conditions), so this correctly handles exploding blocking rules and computes the marginal (additional) and cumulative number of comparisons generated by each rule.

A single record is returned per blocking rule (so a single rule yields a one-element list). When multiple rules are provided, marginal_comparison_count is the marginal number of comparisons generated by that rule (excluding pairs already generated by preceding rules) and cumulative_comparison_count is the running total.

Defaults to 0.05, meaning Splink estimates the count from a deterministic 5% sample of input records on each side of the blocking join. Set record_sample_proportion=1.0 to compute exact counts.

Parameters:

Name Type Description Default
splink_dataframe_or_dataframes SplinkDataFrame | Sequence[SplinkDataFrame]

Input data

required
blocking_rules BlockingRuleLike | Iterable[BlockingRuleLike]

A single blocking rule or an iterable of blocking rules to analyse.

required
link_type user_input_link_type_options

The link type - "link_only", "dedupe_only" or "link_and_dedupe"

required
unique_id_column_name str

Defaults to "unique_id".

'unique_id'
source_dataset_column_name Optional[str]

Defaults to None.

None
record_sample_proportion float

The sampling proportion applied to each side of the blocking join. Defaults to 0.05, meaning Splink estimates the count from a deterministic 5% sample of input records on each side of the blocking join. Set record_sample_proportion=1.0 to compute exact counts.

0.05

Returns:

Type Description
list[CumulativeComparisonRecord]

list[CumulativeComparisonRecord]: One record per blocking rule.

n_largest_blocks(splink_dataframe_or_dataframes, *, blocking_rule, link_type, n_largest=5)

Find the values responsible for creating the largest blocks of records.

For example, when blocking on first name and surname, the 'John Smith' block might be the largest block of records. In cases where values are highly skewed a few values may be resonsible for generating a large proportion of all comparisons. This function helps you find the culprit values.

The analysis is performed pre filter conditions, read more about what this means here

Parameters:

Name Type Description Default
splink_dataframe_or_dataframes SplinkDataFrame | Sequence[SplinkDataFrame]

Input data

required
blocking_rule Union[BlockingRuleCreator, str, Dict[str, Any]]

The blocking rule to analyse

required
link_type user_input_link_type_options

The link type - "link_only", "dedupe_only" or "link_and_dedupe"

required
n_largest int

How many rows to return. Defaults to 5.

5

Returns:

Name Type Description
SplinkDataFrame 'SplinkDataFrame'

A dataframe containing the n_largest blocks