Most “self-improving” agent-skill pipelines I’ve seen in production hit the same wall: after one or two rounds of self-reflection, the delta collapses to noise. The SkillEvo paper names the cause precisely — the evolution gradient decays because the evaluator only speaks in single-turn Q&A, so any defect that requires a follow-up to surface stays invisible forever. That framing alone is worth the read.

Their fix has two pieces I actually care about. First, they re-cast multi-turn user simulation from an evaluation endpoint into a feedback generator: each follow-up question exposes a new layer of defect, so every revision round both consumes and produces gradient. Second, they replace the scalar accept/reject gate with an independent governance layer that actively repairs factual degradation and structural bloat, instead of just rejecting the degraded candidate and losing the signal. Across nine production Skills and 98 skill-reference files, they report +23.0 points over self-reflection loops and +15.4 over single-turn-QA loops.

If you’re operating an agent-skill catalog and your improvement loop plateaus after the first regeneration pass, this is the diagnostic to steal even if you don’t adopt the whole scheme. The HF paper page has the discussion link and reference list. My guess: the “trustworthy feedback keeps supplying gradient” framing is going to reshape how enterprise skill-governance pipelines are wired within a year — the scalar-gate design was always leaking information, we just didn’t name it.

tags: [ agentic-ai ] [ llm-ops ] [ research ]