Reasoned Rule Mining#

Module contents#

Reasoned Rule Mining.

An interpretable binary classifier that mines natural-language IF-THEN rules from LLM reasoning over labelled data, compiles them into a decision policy, and predicts via a calibrated, confidence-weighted ensemble vote with an optional harsh re-evaluation stream.

pydantic model ExtractedRule#

Bases: BaseModel

One extracted IF-THEN rule with its forced outcome and perplexity.

rule#

The natural-language IF ... THEN label = YES/NO rule text.

outcome#

The label the rule is aligned to (the sample’s true label).

perplexity#

Model perplexity of the rule text (lower = more confident).

field outcome: Literal['YES', 'NO'] [Required]#
field perplexity: float [Required]#
field rule: str [Required]#
class RRMConfig(perplexity_threshold=1.6, ensemble_size=3, precision_voting_threshold=0.5, ensemble_confidence_threshold=0.8, final_confidence_threshold=0.6, enable_final_confidence_check=False, use_harsh=True, harsh_level='strict', beta=0.5, recall_floors=(0.05, 0.1, 0.15), weights_alpha_range=(0.0, 5.0), weights_beta_range=(0.0, 5.0), weights_bias_range=(-10.0, 10.0), weights_n_points=11, weights_thresh_points=100, use_memory=True, memory_update_interval=50, random_state=42)#

Bases: object

Configuration for Reasoned Rule Mining.

Defaults mirror the standalone VCBench reference implementation.

Parameters:
  • perplexity_threshold (float) – Keep only rules with perplexity <= this value when compiling the decision policy.

  • ensemble_size (int) – Number of votes per sample in the ensemble (and harsh) stage.

  • precision_voting_threshold (float) – Fraction of votes that must be YES for the ensemble to keep a YES prediction.

  • ensemble_confidence_threshold (float) – A tentative YES is downgraded to NO if the ensemble probability is below this value.

  • final_confidence_threshold (float) – Extra confidence floor applied only when enable_final_confidence_check is True.

  • enable_final_confidence_check (bool) – Whether to apply final_confidence_threshold.

  • use_harsh (bool) – Whether to run the two-stream harsh re-evaluation stage.

  • harsh_level (Literal['light', 'moderate', 'strict']) – Severity of harsh re-evaluation (“light”/”moderate”/”strict”).

  • beta (float) – Beta for the F-beta objective optimised by the combiner (0.5 = F0.5).

  • recall_floors (Tuple[float, ...]) – Recall floors swept during weight optimisation; the floor with the best F-beta is selected.

  • weights_alpha_range (Tuple[float, float]) – (min, max) grid range for the ensemble-stream weight.

  • weights_beta_range (Tuple[float, float]) – (min, max) grid range for the harsh-stream weight.

  • weights_bias_range (Tuple[float, float]) – (min, max) grid range for the combiner bias.

  • weights_n_points (int) – Grid resolution per weight axis.

  • weights_thresh_points (int) – Number of decision thresholds swept per weight tuple.

  • use_memory (bool) – Whether to maintain a rolling reasoning-memory summary.

  • memory_update_interval (int) – Number of samples between memory-summary updates.

  • random_state (int) – Random seed for Platt-scaling logistic regressions.

beta: float = 0.5#
enable_final_confidence_check: bool = False#
ensemble_confidence_threshold: float = 0.8#
ensemble_size: int = 3#
final_confidence_threshold: float = 0.6#
harsh_level: Literal['light', 'moderate', 'strict'] = 'strict'#
memory_update_interval: int = 50#
perplexity_threshold: float = 1.6#
precision_voting_threshold: float = 0.5#
random_state: int = 42#
recall_floors: Tuple[float, ...] = (0.05, 0.1, 0.15)#
use_harsh: bool = True#
use_memory: bool = True#
weights_alpha_range: Tuple[float, float] = (0.0, 5.0)#
weights_beta_range: Tuple[float, float] = (0.0, 5.0)#
weights_bias_range: Tuple[float, float] = (-10.0, 10.0)#
weights_n_points: int = 11#
weights_thresh_points: int = 100#
class ReasonedRuleMining(reason_llmc, extract_llmc=None, predict_llmc=None, config=None, reason_temperature=1.0, predict_temperature=0.0, llm_semaphore_limit=3, save_path=None, name=None, _llm=None)#

Bases: object

Interpretable rule-mining binary classifier.

Parameters:
classmethod load(dir_path)#

Load a model previously saved with save().

Parameters:

dir_path (str | PathLike[str]) – Directory containing reasoned_rule_mining.json.

Return type:

ReasonedRuleMining

Returns:

The restored instance.

Raises:

FileNotFoundError – If the manifest is missing.

async fit(X=None, y=None, *, copy_data=True, reset=False)#

Fit the model: mine rules, compile a policy, and tune the combiner.

Parameters:
  • X (DataFrame | None) – Training features (a DataFrame; each row is stringified).

  • y (Optional[Sequence[str]]) – Training labels, each “YES” or “NO” (case-insensitive).

  • copy_data (bool) – Whether to copy input data.

  • reset (bool) – Clear learned state and refit from scratch.

Return type:

Self

Returns:

The fitted instance.

Raises:
  • ValueError – If set_task has not been called, or data is missing.

  • DataError – If the data is invalid.

async predict(samples)#

Predict labels for samples.

Parameters:

samples (DataFrame) – Samples to predict (a DataFrame; each row is stringified).

Yields:

Tuples of (sample_index, scores, label, token_counter) where scores holds the ensemble/harsh probabilities and the fused combiner score, and label is “YES” or “NO”.

Raises:
  • RuntimeError – If the model has not been fitted.

  • DataError – If samples is not a non-empty DataFrame.

Return type:

AsyncGenerator[Tuple[Any, Dict[str, float], Literal['YES', 'NO'], TokenCounter], None]

save(dir_path=None)#

Persist the fitted model to a directory.

Writes reasoned_rule_mining.json (config + learned state) plus ens_platt.joblib / mod_platt.joblib for the Platt scalers.

Parameters:

dir_path (str | PathLike[str] | None) – Target directory. Defaults to self.save_path.

Raises:

ValueError – If dir_path is an existing file.

Return type:

None

async set_task(task_description)#

Set the binary-classification task description.

The description is interpolated into every prompt, so it should describe what YES and NO mean for the task.

Parameters:

task_description (str) – A description of the binary classification task.

Raises:

ValueError – If task_description is empty.

Return type:

None

property llm_semaphore_limit: int#

Max concurrent LLM calls.

property policy: str#

The compiled decision policy.

Raises:

ValueError – If the model has not been fitted.

property rules: List[ExtractedRule]#

The extracted IF-THEN rules (all, before perplexity filtering).

property task_description: str | None#

The configured task description.

property threshold: float#

The decision threshold on the combiner score.

Raises:

ValueError – If the model has not been fitted.

property token_usage: TokenCounter#

Token counter accumulating usage across all LLM calls.

property validation_result: dict#

Diagnostics from weight optimisation.

Raises:

ValueError – If the model has not been fitted.

property weights: Tuple[float, float, float]#

The fitted combiner weights (alpha, beta, bias).

Raises:

ValueError – If the model has not been fitted.

pydantic model Vote#

Bases: BaseModel

A single YES/NO vote from the prediction LLM.

field vote: Literal['YES', 'NO'] [Required]#