Documentation forsplink.blocking_analysis¶
chart_comparisons_from_blocking_rules(splink_dataframe_or_dataframes, *, blocking_rules, link_type, unique_id_column_name='unique_id', source_dataset_column_name=None, record_sample_proportion=0.05)
¶
Produce a chart of the cumulative number of comparisons generated by one or more blocking rules.
See count_comparisons_from_blocking_rules for details of the underlying
computation and the meaning of record_sample_proportion.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
splink_dataframe_or_dataframes
|
SplinkDataFrame | Sequence[SplinkDataFrame]
|
Input data |
required |
blocking_rules
|
BlockingRuleLike | Iterable[BlockingRuleLike]
|
A single blocking rule or an iterable of blocking rules to analyse. |
required |
link_type
|
user_input_link_type_options
|
The link type - "link_only", "dedupe_only" or "link_and_dedupe" |
required |
unique_id_column_name
|
str
|
Defaults to "unique_id". |
'unique_id'
|
source_dataset_column_name
|
Optional[str]
|
Defaults to None. |
None
|
record_sample_proportion
|
float
|
Defaults to |
0.05
|
Returns:
| Name | Type | Description |
|---|---|---|
CumulativeBlockingRuleComparisonsGeneratedChart |
CumulativeBlockingRuleComparisonsGeneratedChart
|
The chart. |
count_comparisons_from_blocking_rules(splink_dataframe_or_dataframes, *, blocking_rules, link_type, unique_id_column_name='unique_id', source_dataset_column_name=None, record_sample_proportion=0.05)
¶
Analyse one or more blocking rules to understand how many comparisons they will generate.
The comparisons are counted by actually creating the blocked pairs (post filter conditions), so this correctly handles exploding blocking rules and computes the marginal (additional) and cumulative number of comparisons generated by each rule.
A single record is returned per blocking rule (so a single rule yields a
one-element list). When multiple rules are provided,
marginal_comparison_count is the marginal number of comparisons generated by
that rule (excluding pairs already generated by preceding rules) and
cumulative_comparison_count is the running total.
Defaults to 0.05, meaning Splink estimates the count from a deterministic
5% sample of input records on each side of the blocking join. Set
record_sample_proportion=1.0 to compute exact counts.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
splink_dataframe_or_dataframes
|
SplinkDataFrame | Sequence[SplinkDataFrame]
|
Input data |
required |
blocking_rules
|
BlockingRuleLike | Iterable[BlockingRuleLike]
|
A single blocking rule or an iterable of blocking rules to analyse. |
required |
link_type
|
user_input_link_type_options
|
The link type - "link_only", "dedupe_only" or "link_and_dedupe" |
required |
unique_id_column_name
|
str
|
Defaults to "unique_id". |
'unique_id'
|
source_dataset_column_name
|
Optional[str]
|
Defaults to None. |
None
|
record_sample_proportion
|
float
|
The sampling proportion applied
to each side of the blocking join. Defaults to |
0.05
|
Returns:
| Type | Description |
|---|---|
list[CumulativeComparisonRecord]
|
list[CumulativeComparisonRecord]: One record per blocking rule. |
n_largest_blocks(splink_dataframe_or_dataframes, *, blocking_rule, link_type, n_largest=5)
¶
Find the values responsible for creating the largest blocks of records.
For example, when blocking on first name and surname, the 'John Smith' block might be the largest block of records. In cases where values are highly skewed a few values may be resonsible for generating a large proportion of all comparisons. This function helps you find the culprit values.
The analysis is performed pre filter conditions, read more about what this means here
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
splink_dataframe_or_dataframes
|
SplinkDataFrame | Sequence[SplinkDataFrame]
|
Input data |
required |
blocking_rule
|
Union[BlockingRuleCreator, str, Dict[str, Any]]
|
The blocking rule to analyse |
required |
link_type
|
user_input_link_type_options
|
The link type - "link_only", "dedupe_only" or "link_and_dedupe" |
required |
n_largest
|
int
|
How many rows to return. Defaults to 5. |
5
|
Returns:
| Name | Type | Description |
|---|---|---|
SplinkDataFrame |
'SplinkDataFrame'
|
A dataframe containing the n_largest blocks |