Autoresearch feature discovery

Propose Jev questions that turn tasting notes into numeric features for CatBoost.

Research
Result
Held-out RMSE 1.77 after five rounds, versus 3.09 for the mean baseline

CatBoost needs numbers. A tasting note is not a number. This Jev research use case builds the table from questions nobody wrote by hand.

Round 1 proposed 18 questions and reached 1.87 RMSE. Four more rounds of reading the worst predictions reached 1.77 on 800 held-out reviews.

The strongest feature was overall tone positivity, not a chemistry question. Change PROPOSER_TASK to point the same loop at your own labelled text.

Pipeline

  1. An LLM proposes add, revise, or drop actions
  2. Jev answers every new question for all 2,000 notes
  3. CatBoost trains on the encoded answers
  4. Worst and best predictions feed the next proposal round

Builder: TypeSafe cookbook. Stack: jev, claude, catboost.