Montréal
Coucher de soleil
Nuageux · 25°
David Paquet Pitts
/
Radar
On my radar
Week of Jun 8, 2026
8
finds ·
tools
·
libs
·
research
Tools
4 finds
JUN 12
karpathy/autoresearch
Karpathy's experiment in automating research loops with LLMs. Worth a look for harness-design ideas.
Tool
Watch Later
JUN 12
lm-evaluation-harness
EleutherAI's de-facto standard harness for benchmarking LLMs across hundreds of tasks. The reference tool when you need reproducible eval numbers.
Tool
To Test
JUN 11
GEPA (reference implementation)
The official MIT-licensed Python impl of the GEPA optimizer — evolve prompts, code, and arbitrary text artifacts against your own eval.
Tool
To Test
JUN 11
DeepEval
Popular open-source LLM eval framework — pytest-style assertions, RAG/agent metrics, and built-in GEPA-style prompt optimization.
Tool
To Test
Libs
3 finds
JUN 12
inkeep/agents
Open-source framework for building and running agents. Relevant reference point for agent-harness architecture.
Lib
To Test
JUN 11
DSPy — GEPA optimizer
DSPy's first-class GEPA integration. The most ergonomic way to run GEPA over a declarative LLM program with a defined metric.
Lib
To Test
JUN 11
LLMLingua
Microsoft's prompt-compression toolkit — compresses prompts up to 20× with minimal performance loss. The serious counterpoint to anecdotal SPR-style compression.
Lib
To Test
Research
1 find
JUN 11
GEPA
Reflective Prompt Evolution Can Outperform RL
Research
Wow
← Week of May 25, 2026
Week of Jun 15, 2026 →