Reasoned Rule Mining#
Module contents#
Reasoned Rule Mining.
An interpretable binary classifier that mines natural-language IF-THEN rules from LLM reasoning over labelled data, compiles them into a decision policy, and predicts via a calibrated, confidence-weighted ensemble vote with an optional harsh re-evaluation stream.
- pydantic model ExtractedRule#
Bases:
BaseModelOne extracted IF-THEN rule with its forced outcome and perplexity.
- rule#
The natural-language
IF ... THEN label = YES/NOrule text.
- outcome#
The label the rule is aligned to (the sample’s true label).
- perplexity#
Model perplexity of the rule text (lower = more confident).
- field outcome: Literal['YES', 'NO'] [Required]#
- field perplexity: float [Required]#
- field rule: str [Required]#
- class RRMConfig(perplexity_threshold=1.6, ensemble_size=3, precision_voting_threshold=0.5, ensemble_confidence_threshold=0.8, final_confidence_threshold=0.6, enable_final_confidence_check=False, use_harsh=True, harsh_level='strict', beta=0.5, recall_floors=(0.05, 0.1, 0.15), weights_alpha_range=(0.0, 5.0), weights_beta_range=(0.0, 5.0), weights_bias_range=(-10.0, 10.0), weights_n_points=11, weights_thresh_points=100, use_memory=True, memory_update_interval=50, random_state=42)#
Bases:
objectConfiguration for Reasoned Rule Mining.
Defaults mirror the standalone VCBench reference implementation.
- Parameters:
perplexity_threshold (
float) – Keep only rules with perplexity <= this value when compiling the decision policy.ensemble_size (
int) – Number of votes per sample in the ensemble (and harsh) stage.precision_voting_threshold (
float) – Fraction of votes that must be YES for the ensemble to keep a YES prediction.ensemble_confidence_threshold (
float) – A tentative YES is downgraded to NO if the ensemble probability is below this value.final_confidence_threshold (
float) – Extra confidence floor applied only whenenable_final_confidence_checkis True.enable_final_confidence_check (
bool) – Whether to applyfinal_confidence_threshold.use_harsh (
bool) – Whether to run the two-stream harsh re-evaluation stage.harsh_level (
Literal['light','moderate','strict']) – Severity of harsh re-evaluation (“light”/”moderate”/”strict”).beta (
float) – Beta for the F-beta objective optimised by the combiner (0.5 = F0.5).recall_floors (
Tuple[float,...]) – Recall floors swept during weight optimisation; the floor with the best F-beta is selected.weights_alpha_range (
Tuple[float,float]) – (min, max) grid range for the ensemble-stream weight.weights_beta_range (
Tuple[float,float]) – (min, max) grid range for the harsh-stream weight.weights_bias_range (
Tuple[float,float]) – (min, max) grid range for the combiner bias.weights_n_points (
int) – Grid resolution per weight axis.weights_thresh_points (
int) – Number of decision thresholds swept per weight tuple.use_memory (
bool) – Whether to maintain a rolling reasoning-memory summary.memory_update_interval (
int) – Number of samples between memory-summary updates.random_state (
int) – Random seed for Platt-scaling logistic regressions.
- class ReasonedRuleMining(reason_llmc, extract_llmc=None, predict_llmc=None, config=None, reason_temperature=1.0, predict_temperature=0.0, llm_semaphore_limit=3, save_path=None, name=None, _llm=None)#
Bases:
objectInterpretable rule-mining binary classifier.
- Parameters:
reason_llmc (
List[Union[AnthropicChoice,GoogleChoice,OpenAIChoice,XAIChoice,AnthropicChoiceDict,GoogleChoiceDict,OpenAIChoiceDict,XAIChoiceDict]]) – LLMs for reasoning, policy compilation, and memory summarisation, in priority order.extract_llmc (
Optional[List[Union[AnthropicChoice,GoogleChoice,OpenAIChoice,XAIChoice,AnthropicChoiceDict,GoogleChoiceDict,OpenAIChoiceDict,XAIChoiceDict]]]) – LLMs for rule extraction. Defaults toreason_llmc.predict_llmc (
Optional[List[Union[AnthropicChoice,GoogleChoice,OpenAIChoice,XAIChoice,AnthropicChoiceDict,GoogleChoiceDict,OpenAIChoiceDict,XAIChoiceDict]]]) – LLMs for ensemble voting and harsh re-evaluation. Defaults toreason_llmc.config (
RRMConfig|None) – Algorithm configuration. Defaults toRRMConfig().reason_temperature (
float) – Sampling temperature for reasoning/policy/memory.predict_temperature (
float) – Sampling temperature for voting/harsh stages.llm_semaphore_limit (
int) – Max concurrent LLM calls.save_path (
str|PathLike[str] |None) – Directory for checkpoints/models.name (
str|None) – Instance name (alphanumeric/underscore)._llm (
Any) – LLM instance for dependency injection (testing). If None, uses the globalllmsingleton.
- classmethod load(dir_path)#
Load a model previously saved with
save().- Parameters:
dir_path (
str|PathLike[str]) – Directory containingreasoned_rule_mining.json.- Return type:
- Returns:
The restored instance.
- Raises:
FileNotFoundError – If the manifest is missing.
- async fit(X=None, y=None, *, copy_data=True, reset=False)#
Fit the model: mine rules, compile a policy, and tune the combiner.
- Parameters:
- Return type:
Self- Returns:
The fitted instance.
- Raises:
ValueError – If
set_taskhas not been called, or data is missing.DataError – If the data is invalid.
- async predict(samples)#
Predict labels for samples.
- Parameters:
samples (
DataFrame) – Samples to predict (a DataFrame; each row is stringified).- Yields:
Tuples of
(sample_index, scores, label, token_counter)wherescoresholds the ensemble/harsh probabilities and the fused combinerscore, andlabelis “YES” or “NO”.- Raises:
RuntimeError – If the model has not been fitted.
DataError – If
samplesis not a non-empty DataFrame.
- Return type:
AsyncGenerator[Tuple[Any,Dict[str,float],Literal['YES','NO'],TokenCounter],None]
- save(dir_path=None)#
Persist the fitted model to a directory.
Writes
reasoned_rule_mining.json(config + learned state) plusens_platt.joblib/mod_platt.joblibfor the Platt scalers.
- async set_task(task_description)#
Set the binary-classification task description.
The description is interpolated into every prompt, so it should describe what YES and NO mean for the task.
- Parameters:
task_description (
str) – A description of the binary classification task.- Raises:
ValueError – If
task_descriptionis empty.- Return type:
- property policy: str#
The compiled decision policy.
- Raises:
ValueError – If the model has not been fitted.
- property rules: List[ExtractedRule]#
The extracted IF-THEN rules (all, before perplexity filtering).
- property threshold: float#
The decision threshold on the combiner score.
- Raises:
ValueError – If the model has not been fitted.
- property token_usage: TokenCounter#
Token counter accumulating usage across all LLM calls.
- property validation_result: dict#
Diagnostics from weight optimisation.
- Raises:
ValueError – If the model has not been fitted.