Feature Generation#

Feature Generation asks a language model to write binary rules as small Python lambda expressions, then runs those rules on your records to build a 0 or 1 feature matrix. The model is called once to write the rules. Evaluating them afterwards costs nothing, is deterministic, and every feature is a line of code a person can read.

How it works#

  1. Describe your records with a DataSchema: what a record is, its fields, the name of the lambda parameter and a few example rules in the style you want.

  2. FeatureGenerator shows the model a sample of labelled records and asks for rules, each with a name, a description and a lambda expression.

  3. FeatureEvaluator compiles the expressions and applies them to any number of records, returning one column per rule.

The resulting columns can go into any classifier, or be read directly as the reasons behind a prediction.

Example#

Run it in Jupyter, which allows await at the top level, with OPENAI_API_KEY set.

from think_reason_learn.core.llms import OpenAIChoice
from think_reason_learn.features import (
    DataSchema, FeatureEvaluator, FeatureGenerator, HelperFunction, Rule,
)

founders = [
    {"degrees": ["BSc", "PhD"], "prior_exits": 2, "years_experience": 12},
    {"degrees": ["BA"], "prior_exits": 0, "years_experience": 3},
    {"degrees": ["MBA"], "prior_exits": 1, "years_experience": 9},
    {"degrees": [], "prior_exits": 0, "years_experience": 1},
]
labels = [1, 0, 1, 0]

schema = DataSchema(
    description="An anonymised founder profile.",
    schema_text="founder = {'degrees': list[str], 'prior_exits': int, 'years_experience': int}",
    param_name="founder",
    example_rules=[
        Rule(
            name="has_prior_exit",
            description="The founder has sold or taken public a company before.",
            expression="lambda founder: founder.get('prior_exits', 0) > 0",
        ),
    ],
)

# Optional: functions the model may call inside its rules
has_doctorate = HelperFunction(
    name="has_doctorate",
    func=lambda degrees: any(d.lower() == "phd" for d in degrees),
    signature="has_doctorate(degrees: list[str]) -> bool",
    docstring="True when one of the degrees is a PhD.",
)

generator = FeatureGenerator(
    schema=schema,
    helpers=[has_doctorate],
    llm_priority=[OpenAIChoice(model="gpt-4o-mini")],
)
rules = await generator.generate(founders, labels, n_rules=10)

evaluator = FeatureEvaluator(rules, helpers=[has_doctorate])
features = evaluator.evaluate_df(founders)  # one 0 or 1 column per rule
print(evaluator.compilation_errors)          # rules that did not compile, if any

The evaluator works without a model too, so you can check or hand-write rules:

evaluator = FeatureEvaluator(schema.example_rules)
evaluator.evaluate(founders[0])  # {'has_prior_exit': 1}

Improving the rules#

Pass the rules from one round back in with prior_rules and the model uses them as feedback for the next set:

more_rules = await generator.generate(founders, labels, n_rules=10, prior_rules=rules)

n_samples sets how many labelled records go into the prompt (60 by default), and temperature on the generator sets how varied the rules are. Pass temperature=None for reasoning models that do not accept one.

Cognitive modes#

cognitive_modes adds structured reasoning instructions to the prompt, following the CoFEE framework (Westermann, 2025). Choose any of CognitiveMode: BACKWARD_CHAINING, SUBGOAL_DECOMPOSITION, VERIFICATION and BACKTRACKING. Without it the prompt is unchanged.

from think_reason_learn.features import CognitiveMode

generator = FeatureGenerator(
    schema=schema,
    helpers=[has_doctorate],
    llm_priority=[OpenAIChoice(model="gpt-4o-mini")],
    cognitive_modes={CognitiveMode.BACKWARD_CHAINING, CognitiveMode.VERIFICATION},
)

Safety#

Generated expressions run with a restricted set of built-ins: len, any, all, sum, min, max, abs, sorted and similar helpers, the basic type constructors and isinstance, plus the helper functions you pass in. __import__, eval, exec, open, getattr and the other built-ins that reach outside the record are not available. A rule that fails to compile, or raises on a record, scores 0 for that record, and compilation_errors lists the rules that did not compile.

This is a guard against mistakes in generated code, not a sandbox for code you do not trust.

Module contents#

Lambda-based feature generation.

Generate binary classification rules as executable Python lambda expressions via an LLM, then evaluate them deterministically on structured data to produce a binary feature matrix — at zero marginal cost per evaluation.

class CognitiveMode(*values)#

Bases: StrEnum

Cognitive reasoning behaviours that can be toggled in the generation prompt.

Based on the CoFEE framework (Westermann, 2025), which enforces structured cognitive behaviours during LLM-based feature discovery. Each mode injects a corresponding section into the system prompt.

BACKTRACKING = 'backtracking'#
BACKWARD_CHAINING = 'backward_chaining'#
SUBGOAL_DECOMPOSITION = 'subgoal_decomposition'#
VERIFICATION = 'verification'#
class DataSchema(description, schema_text, param_name, example_rules)#

Bases: object

Description of the input data structure for the LLM prompt.

Parameters:
description#

Plain-English description of the record type.

schema_text#

A Python-like schema definition inserted verbatim into the system prompt (e.g. the dict structure with types).

param_name#

The lambda parameter name (e.g. "founder").

example_rules#

A few hand-written Rule instances that demonstrate the expected lambda style.

description: str#
example_rules: list[Rule]#
param_name: str#
schema_text: str#
class FeatureEvaluator(rules, helpers=None)#

Bases: object

Compile and evaluate LLM-generated lambda rules on structured records.

Parameters:
  • rules (Sequence[Rule]) – Rules whose expression fields are Python lambda strings.

  • helpers (Optional[Sequence[HelperFunction]]) – Optional helper functions to make available inside the lambdas (e.g. parse_qs, parse_duration). Each helper’s .func is placed into the eval context under its .name.

evaluate(record)#

Evaluate all rules on a single record.

Returns a dict mapping rule_name → 0 | 1. Rules that fail to compile or that raise at runtime silently return 0.

Return type:

dict[str, int]

Parameters:

record (dict[str, Any])

evaluate_df(records)#

Evaluate all rules on multiple records.

Returns a DataFrame with one column per rule (values 0 or 1) and one row per record.

Return type:

DataFrame

Parameters:

records (Sequence[dict[str, Any]])

property compilation_errors: dict[str, str]#

Rules that failed to compile, keyed by rule name.

property rules: list[Rule]#

The rules this evaluator was initialised with.

class FeatureGenerator(schema, helpers, llm_priority, temperature=0.7, cognitive_modes=None)#

Bases: object

Generate binary classification rules via an LLM.

Parameters:
async generate(samples, labels, *, n_rules=30, n_samples=60, prior_rules=None)#

Generate binary classification rules from structured data samples.

Parameters:
  • samples (Sequence[dict[str, Any]]) – All available structured data records (dicts).

  • labels (Sequence[Any]) – Parallel label sequence (one per record, e.g. 1 / 0).

  • n_rules (int) – Number of rules to request from the LLM.

  • n_samples (int) – Number of sample records to include in the prompt.

  • prior_rules (Optional[Sequence[Rule]]) – Optional rules from a previous iteration, injected as feedback.

  • Returns

  • -------

  • list[Rule] – The generated rules, ready for FeatureEvaluator.

Return type:

list[Rule]

pydantic model GeneratedRule#

Bases: BaseModel

A single binary classification rule returned by the LLM.

field description: str [Required]#

Short description of what this rule checks.

field expression: str [Required]#

A Python lambda expression that takes one dict argument and returns True or False.

field name: str [Required]#

Snake_case identifier for the rule, e.g. ‘has_phd’ or ‘prior_exit’.

pydantic model GeneratedRules#

Bases: BaseModel

Structured output containing all generated rules.

field rules: List[GeneratedRule] [Required]#

List of generated binary classification rules.

class HelperFunction(name, func, signature, docstring)#

Bases: object

A helper function available to lambda expressions.

The function must exist in both worlds: - In the LLM prompt (so the model knows it can use it) - In the eval context (so the lambda can call it at runtime)

Parameters:
docstring: str#
func: Callable[..., Any]#
name: str#
signature: str#
class Rule(name, description, expression)#

Bases: object

A compiled-ready rule with name, description, and expression string.

Parameters:
  • name (str)

  • description (str)

  • expression (str)

description: str#
expression: str#
name: str#