Index A | B | C | D | E | F | G | H | I | L | M | N | O | P | Q | R | S | T | U | V | W | X A action_dim (VerifiableRL property) ActionSupervisor (class in think_reason_learn.verifiable_rl) add_question() (RRF method) advice() (GPTree method) aggregate_predictions() (RRF static method) AnsSimilarityFunc (in module think_reason_learn.rrf._types) answer (Answer attribute) ANTHROPIC_NOT_GIVEN (NotGiven attribute) AnthropicChoiceDict (class in think_reason_learn.core.llms) append() (TokenCounter method) aprefer() (LLMActionSupervisor method) average_confidence (LLMResponse property) B beta (RRMConfig attribute) (WeightTrainerConfig attribute) C callers (TokenCount attribute) children (Node attribute) choices (NodeQuestion attribute) class_distribution (Node attribute) class_weight_balanced (WeightTrainerConfig attribute) classes (GPTree property) clf_batch (VerifiableRLConfig attribute) clf_hidden (VerifiableRLConfig attribute) clf_lr (VerifiableRLConfig attribute) clf_replay_max (VerifiableRLConfig attribute) clf_sample (VerifiableRLConfig attribute) clf_target_update_every (VerifiableRLConfig attribute) clf_threshold (VerifiableRLConfig attribute) COST_PRUNING (QuestionExclusion attribute) CostSensitiveConfig (class in think_reason_learn.rrf) Criterion (in module think_reason_learn.gptree._types) critic_instructions_template (GPTree property) cross_validate_aggregation() (in module think_reason_learn.rrf) Cs (WeightTrainerConfig attribute) cumulative_memory (Node attribute) cv_folds (WeightTrainerConfig attribute) CVResult (class in think_reason_learn.rrf) D decision_path (QueryResult attribute), [1] description (PromptPreset attribute), [1] df_column (NodeQuestion attribute) E EmbeddingModel (in module think_reason_learn.rrf._types) enable_final_confidence_check (RRMConfig attribute) enable_semantic_filter (CostSensitiveConfig attribute) ensemble_confidence_threshold (RRMConfig attribute) ensemble_size (RRMConfig attribute) eps (VerifiableRLConfig attribute) exclusion_report() (RRF method) EXPERT (QuestionExclusion attribute) expert_advice (GPTree property) F filter_questions_on_pred_similarity() (RRF method) filter_questions_on_semantics() (RRF method) final_confidence_threshold (RRMConfig attribute) fit() (GPTree method) (PolicyInduction method) (ReasonedRuleMining method) (RRF method) (VerifiableRL method) fold_metrics (CVResult attribute), [1] freeze_clf_updates (VerifiableRLConfig attribute) from_dict() (Node class method) (NodeQuestion class method) (TokenCount class method) (TokenCounter class method) from_state_dicts() (VerifiableRL class method) G get_answers() (RRF method) get_leaf_prediction() (GPTree method) get_leaf_proba() (GPTree method) get_memory() (PolicyInduction method) get_node() (GPTree method) get_questions() (GPTree method) (RRF method) get_root_id() (GPTree method) get_training_data() (GPTree method) gini (Node attribute) GoogleChoiceDict (class in think_reason_learn.core.llms) GPTree (class in think_reason_learn.gptree) grad_clip (VerifiableRLConfig attribute) greedy (VerifiableRLConfig attribute) H harsh_level (RRMConfig attribute) I id (Node attribute) is_fitted (VerifiableRL property) is_leaf (Node property) is_min_estimate (TokenCount attribute) L label (Node attribute) LLM (class in think_reason_learn.core.llms) llm (in module think_reason_learn.core.llms) llm_bias (VerifiableRLConfig attribute) llm_semaphore_limit (PolicyInduction property) (ReasonedRuleMining property) (RRF property) LLMActionSupervisor (class in think_reason_learn.verifiable_rl) LLMChoice (in module think_reason_learn.core.llms._schemas) LLMChoiceDict (in module think_reason_learn.core.llms._schemas) LLMChoiceModel (in module think_reason_learn.core.llms._schemas) load() (GPTree class method) (PolicyInduction class method) (ReasonedRuleMining class method) (RRF class method) (VerifiableRL class method) logprobs (LLMResponse attribute) lr (PolicyInduction property) M max_depth (VerifiableRLConfig attribute) max_questions_full_eval (CostSensitiveConfig attribute) max_screening_samples (CostSensitiveConfig attribute) max_steps (VerifiableRLConfig attribute) memory_update_interval (RRMConfig attribute) min_queries (VerifiableRLConfig attribute) model (AnthropicChoice attribute) (AnthropicChoiceDict attribute) (GoogleChoice attribute) (GoogleChoiceDict attribute) (OpenAIChoice attribute) (OpenAIChoiceDict attribute) (TokenCount attribute) (XAIChoice attribute) (XAIChoiceDict attribute) module think_reason_learn think_reason_learn.core think_reason_learn.core.llms think_reason_learn.core.utils think_reason_learn.gptree think_reason_learn.policy_induction think_reason_learn.reasoned_rule_mining think_reason_learn.rrf think_reason_learn.verifiable_rl N n_iterations (VerifiableRLConfig attribute) n_rollouts (VerifiableRLConfig attribute) name (PromptPreset attribute), [1] Node (class in think_reason_learn.gptree) NodeQuestion (class in think_reason_learn.gptree) NotGiven (class in think_reason_learn.core.llms) number_of_calls (TokenCount attribute) O OPENAI_NOT_GIVEN (NotGiven attribute) OpenAIChoiceDict (class in think_reason_learn.core.llms) outcome (ExtractedRule attribute), [1] P parent_id (Node attribute) penalty (WeightTrainerConfig attribute) per_founder (CVResult attribute), [1] perplexity (ExtractedRule attribute), [1] perplexity_threshold (RRMConfig attribute) policies (Policies attribute) policy (ReasonedRuleMining property) policy_batch (VerifiableRLConfig attribute) policy_gen_instructions_template (PolicyInduction property) policy_hidden (VerifiableRLConfig attribute) policy_lr (VerifiableRLConfig attribute) policy_replay_max (VerifiableRLConfig attribute) policy_sample (VerifiableRLConfig attribute) PolicyInduction (class in think_reason_learn.policy_induction) precision_voting_threshold (RRMConfig attribute) predict() (GPTree method) (PolicyInduction method) (ReasonedRuleMining method) (RRF method) (VerifiableRL method) predict_founder_level() (RRF method) predict_min_queries (VerifiableRLConfig attribute) predict_paths() (VerifiableRL method) predict_proba() (VerifiableRL method) predict_threshold (VerifiableRLConfig attribute) prediction (QueryResult attribute), [1] PREDICTION_SIMILARITY (QuestionExclusion attribute) prefer (NextActionPreference attribute), [1] prefer() (ActionSupervisor method) (LLMActionSupervisor method) pretrain_classifier() (VerifiableRL method) probability (QueryResult attribute), [1] PromptPreset (class in think_reason_learn.rrf) provider (AnthropicChoice attribute) (AnthropicChoiceDict attribute) (GoogleChoice attribute) (GoogleChoiceDict attribute) (OpenAIChoice attribute) (OpenAIChoiceDict attribute) (TokenCount attribute) (XAIChoice attribute) (XAIChoiceDict attribute) provider_model (LLMResponse attribute) prune_tree() (GPTree method) Q QueryResult (class in think_reason_learn.verifiable_rl) question (Node attribute) question_answer_system (PromptPreset attribute), [1] question_answer_user_template (PromptPreset attribute), [1] question_gen_instructions_template (GPTree property) (RRF property) question_gen_system (PromptPreset attribute), [1] question_gen_user_template (PromptPreset attribute), [1] question_type (NodeQuestion attribute) QuestionExclusion (class in think_reason_learn.rrf) questions (Node attribute) QuestionType (in module think_reason_learn.gptree._types) R random_state (RRMConfig attribute) (WeightTrainerConfig attribute) ReasonedRuleMining (class in think_reason_learn.reasoned_rule_mining) recall_floors (RRMConfig attribute) repeat_penalty (VerifiableRLConfig attribute) respond() (LLM method) respond_sync() (LLM method) response (LLMResponse attribute) resume_fit() (GPTree method) reward_fn (VerifiableRLConfig attribute) reward_fp (VerifiableRLConfig attribute) reward_tn (VerifiableRLConfig attribute) reward_tp (VerifiableRLConfig attribute) RRF (class in think_reason_learn.rrf) RRMConfig (class in think_reason_learn.reasoned_rule_mining) rule (ExtractedRule attribute), [1] rules (ReasonedRuleMining property) S save() (GPTree method) (PolicyInduction method) (ReasonedRuleMining method) (RRF method) (VerifiableRL method) score (NodeQuestion attribute) screening_baseline (CostSensitiveConfig attribute) screening_fraction (CostSensitiveConfig attribute) screening_metric (CostSensitiveConfig attribute) semantic_emb_model (CostSensitiveConfig attribute) semantic_threshold (CostSensitiveConfig attribute) SEMANTICS (QuestionExclusion attribute) set_task() (PolicyInduction method) (ReasonedRuleMining method) set_tasks() (GPTree method) (RRF method) slots_used (QueryResult attribute), [1] split_ratios (Node attribute) state_dim (VerifiableRL property) step_penalty (VerifiableRLConfig attribute) stop() (GPTree method) summary (CVResult attribute), [1] T task_description (GPTree property) (PolicyInduction property) (ReasonedRuleMining property) (RRF property) tau_info (VerifiableRLConfig attribute) tau_stop (VerifiableRLConfig attribute) think_reason_learn module think_reason_learn.core module think_reason_learn.core.llms module think_reason_learn.core.utils module think_reason_learn.gptree module think_reason_learn.policy_induction module think_reason_learn.reasoned_rule_mining module think_reason_learn.rrf module think_reason_learn.verifiable_rl module threshold (PolicyInduction property) (ReasonedRuleMining property) threshold_grid (WeightTrainerConfig attribute) to_dict() (NodeQuestion method) (TokenCount method) (TokenCounter method) token_counts (TokenCounter attribute) token_usage (GPTree property) (LLMActionSupervisor property) (PolicyInduction property) (ReasonedRuleMining property) (RRF property) TokenCount (class in think_reason_learn.core.llms) TokenCounter (class in think_reason_learn.core.llms) total_tokens (LLMResponse attribute) train_epochs (VerifiableRLConfig attribute) U uncertain_delta (VerifiableRLConfig attribute) update_every (VerifiableRLConfig attribute) update_question_exclusion() (RRF method) use_harsh (RRMConfig attribute) use_memory (RRMConfig attribute) V validation_result (PolicyInduction property) (ReasonedRuleMining property) value (NodeQuestion attribute) (TokenCount attribute) VerifiableRL (class in think_reason_learn.verifiable_rl) VerifiableRLConfig (class in think_reason_learn.verifiable_rl) view_node() (GPTree method) vote (Vote attribute) W weights (ReasonedRuleMining property) weights_alpha_range (RRMConfig attribute) weights_beta_range (RRMConfig attribute) weights_bias_range (RRMConfig attribute) weights_n_points (RRMConfig attribute) weights_thresh_points (RRMConfig attribute) WeightTrainerConfig (class in think_reason_learn.policy_induction) X XAIChoiceDict (class in think_reason_learn.core.llms)