Feature Generation#
Feature Generation asks a language model to write binary rules as small Python lambda expressions, then runs those rules on your records to build a 0 or 1 feature matrix. The model is called once to write the rules. Evaluating them afterwards costs nothing, is deterministic, and every feature is a line of code a person can read.
How it works#
Describe your records with a
DataSchema: what a record is, its fields, the name of the lambda parameter and a few example rules in the style you want.FeatureGeneratorshows the model a sample of labelled records and asks for rules, each with a name, a description and a lambda expression.FeatureEvaluatorcompiles the expressions and applies them to any number of records, returning one column per rule.
The resulting columns can go into any classifier, or be read directly as the reasons behind a prediction.
Example#
Run it in Jupyter, which allows await at the top level, with OPENAI_API_KEY set.
from think_reason_learn.core.llms import OpenAIChoice
from think_reason_learn.features import (
DataSchema, FeatureEvaluator, FeatureGenerator, HelperFunction, Rule,
)
founders = [
{"degrees": ["BSc", "PhD"], "prior_exits": 2, "years_experience": 12},
{"degrees": ["BA"], "prior_exits": 0, "years_experience": 3},
{"degrees": ["MBA"], "prior_exits": 1, "years_experience": 9},
{"degrees": [], "prior_exits": 0, "years_experience": 1},
]
labels = [1, 0, 1, 0]
schema = DataSchema(
description="An anonymised founder profile.",
schema_text="founder = {'degrees': list[str], 'prior_exits': int, 'years_experience': int}",
param_name="founder",
example_rules=[
Rule(
name="has_prior_exit",
description="The founder has sold or taken public a company before.",
expression="lambda founder: founder.get('prior_exits', 0) > 0",
),
],
)
# Optional: functions the model may call inside its rules
has_doctorate = HelperFunction(
name="has_doctorate",
func=lambda degrees: any(d.lower() == "phd" for d in degrees),
signature="has_doctorate(degrees: list[str]) -> bool",
docstring="True when one of the degrees is a PhD.",
)
generator = FeatureGenerator(
schema=schema,
helpers=[has_doctorate],
llm_priority=[OpenAIChoice(model="gpt-4o-mini")],
)
rules = await generator.generate(founders, labels, n_rules=10)
evaluator = FeatureEvaluator(rules, helpers=[has_doctorate])
features = evaluator.evaluate_df(founders) # one 0 or 1 column per rule
print(evaluator.compilation_errors) # rules that did not compile, if any
The evaluator works without a model too, so you can check or hand-write rules:
evaluator = FeatureEvaluator(schema.example_rules)
evaluator.evaluate(founders[0]) # {'has_prior_exit': 1}
Improving the rules#
Pass the rules from one round back in with prior_rules and the model uses them as feedback for the next set:
more_rules = await generator.generate(founders, labels, n_rules=10, prior_rules=rules)
n_samples sets how many labelled records go into the prompt (60 by default), and temperature on the generator
sets how varied the rules are. Pass temperature=None for reasoning models that do not accept one.
Cognitive modes#
cognitive_modes adds structured reasoning instructions to the prompt, following the CoFEE framework (Westermann,
2025). Choose any of CognitiveMode: BACKWARD_CHAINING,
SUBGOAL_DECOMPOSITION, VERIFICATION and BACKTRACKING. Without it the prompt is unchanged.
from think_reason_learn.features import CognitiveMode
generator = FeatureGenerator(
schema=schema,
helpers=[has_doctorate],
llm_priority=[OpenAIChoice(model="gpt-4o-mini")],
cognitive_modes={CognitiveMode.BACKWARD_CHAINING, CognitiveMode.VERIFICATION},
)
Safety#
Generated expressions run with a restricted set of built-ins: len, any, all, sum, min, max,
abs, sorted and similar helpers, the basic type constructors and isinstance, plus the helper functions you
pass in. __import__, eval, exec, open, getattr and the other built-ins that reach outside the
record are not available. A rule that fails to compile, or raises on a record, scores 0 for that record, and
compilation_errors lists the rules that did not compile.
This is a guard against mistakes in generated code, not a sandbox for code you do not trust.
Module contents#
Lambda-based feature generation.
Generate binary classification rules as executable Python lambda expressions via an LLM, then evaluate them deterministically on structured data to produce a binary feature matrix — at zero marginal cost per evaluation.
- class CognitiveMode(*values)#
Bases:
StrEnumCognitive reasoning behaviours that can be toggled in the generation prompt.
Based on the CoFEE framework (Westermann, 2025), which enforces structured cognitive behaviours during LLM-based feature discovery. Each mode injects a corresponding section into the system prompt.
- BACKTRACKING = 'backtracking'#
- BACKWARD_CHAINING = 'backward_chaining'#
- SUBGOAL_DECOMPOSITION = 'subgoal_decomposition'#
- VERIFICATION = 'verification'#
- class DataSchema(description, schema_text, param_name, example_rules)#
Bases:
objectDescription of the input data structure for the LLM prompt.
- description#
Plain-English description of the record type.
- schema_text#
A Python-like schema definition inserted verbatim into the system prompt (e.g. the dict structure with types).
- param_name#
The lambda parameter name (e.g.
"founder").
- class FeatureEvaluator(rules, helpers=None)#
Bases:
objectCompile and evaluate LLM-generated lambda rules on structured records.
- Parameters:
rules (
Sequence[Rule]) – Rules whoseexpressionfields are Python lambda strings.helpers (
Optional[Sequence[HelperFunction]]) – Optional helper functions to make available inside the lambdas (e.g.parse_qs,parse_duration). Each helper’s.funcis placed into the eval context under its.name.
- evaluate(record)#
Evaluate all rules on a single record.
Returns a dict mapping
rule_name → 0 | 1. Rules that fail to compile or that raise at runtime silently return0.
- class FeatureGenerator(schema, helpers, llm_priority, temperature=0.7, cognitive_modes=None)#
Bases:
objectGenerate binary classification rules via an LLM.
- Parameters:
schema (
DataSchema) – Describes the input data structure, the lambda parameter name, and includes example rules for the prompt.helpers (
Sequence[HelperFunction]) – Helper functions to advertise in the prompt (the LLM will reference them inside the generated lambda expressions).llm_priority (
Sequence[Union[AnthropicChoice,GoogleChoice,OpenAIChoice,XAIChoice,AnthropicChoiceDict,GoogleChoiceDict,OpenAIChoiceDict,XAIChoiceDict]]) – LLM provider / model choices, tried in order (seeLLMChoice).temperature (
float|None) – Sampling temperature for the LLM call. PassNoneto omit the parameter (required for reasoning models likeo3-mini).cognitive_modes (
set[CognitiveMode] |None) – Optional set ofCognitiveModevalues that inject cognitive reasoning constraints into the system prompt (CoFEE-style). WhenNoneor empty, the prompt is unchanged (baseline behaviour).
- async generate(samples, labels, *, n_rules=30, n_samples=60, prior_rules=None)#
Generate binary classification rules from structured data samples.
- Parameters:
samples (
Sequence[dict[str,Any]]) – All available structured data records (dicts).labels (
Sequence[Any]) – Parallel label sequence (one per record, e.g.1/0).n_rules (
int) – Number of rules to request from the LLM.n_samples (
int) – Number of sample records to include in the prompt.prior_rules (
Optional[Sequence[Rule]]) – Optional rules from a previous iteration, injected as feedback.Returns
-------
list[Rule] – The generated rules, ready for
FeatureEvaluator.
- Return type:
- pydantic model GeneratedRule#
Bases:
BaseModelA single binary classification rule returned by the LLM.
- field description: str [Required]#
Short description of what this rule checks.
- field expression: str [Required]#
A Python lambda expression that takes one dict argument and returns True or False.
- field name: str [Required]#
Snake_case identifier for the rule, e.g. ‘has_phd’ or ‘prior_exit’.
- pydantic model GeneratedRules#
Bases:
BaseModelStructured output containing all generated rules.
- field rules: List[GeneratedRule] [Required]#
List of generated binary classification rules.