By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Microsoft SkillOpt Optimizes AI Agent Skills Across Models
Microsoft researchers have developed SkillOpt, a novel text-space optimizer designed to enhance the transferability of AI agent skills across various model scales and between different AI architectures. The system trains a single natural-language skill document while keeping the target AI model's parameters frozen, a process that significantly reduces computational overhead. An optimizer model analyzes scored outputs from the AI agent and proposes specific edits, such as additions, deletions, or replacements, to the skill document. A validation mechanism ensures that an edit is only accepted if it strictly improves the AI's performance score on a held-out dataset. The final output of this optimization process is a single file, typically named 'best_skill.md', which encapsulates the refined skill.
SkillOpt's effectiveness is demonstrated through transfer tables that report baseline performance, direct optimization scores, and transferred skill scores. The baseline represents the target AI model's performance without any specialized skill. Direct optimization involves training SkillOpt on the specific target model and task. Transferred skill refers to a skill that was trained on a different model or task and then applied to the target model without further in-domain optimization. The critical metric for evaluating SkillOpt is not the absolute performance of the transferred skill, but rather how much of the performance gain achieved through direct optimization is retained when the skill is transferred.
Cross-model transfer experiments within the GPT family, specifically using GPT-5.4 variants, reveal varying degrees of skill retention. For instance, a skill trained on GPT-5.4 and deployed on GPT-5.4-mini for the SpreadsheetBench task retained 82% of the in-domain gain, achieving a score of 47.5 compared to a baseline of 36.1. This high retention suggests a near 'free reuse' of learned procedures. However, performance varied significantly across tasks and model sizes. On the LiveMath task with GPT-5.4-nano, the transferred skill surprisingly outperformed the direct optimization result (28.8 vs. 27.2), which the researchers interpret as evidence of certain learned procedures being agnostic to the specific target model. Conversely, the SpreadsheetBench task on GPT-5.4-nano showed weaker retention at 16% (42.5 direct vs. 26.5 transferred), indicating that skill retention is not uniform. Importantly, all tested rows demonstrated that the transferred skills did not fall below the target model's original no-skill baseline performance.
The research also explores cross-family transfer, such as moving skills from GPT models to Qwen models, indicating the potential for SkillOpt to bridge performance gaps between fundamentally different AI architectures. This capability is crucial for developing more adaptable and efficient AI systems that can leverage pre-trained skills across a wider range of applications and model types, reducing the need for extensive retraining for each new deployment. The development of SkillOpt represents a significant step towards more generalized AI agent capabilities.
Original source — read the full reporting at the publisher:
Read on MarkTechPostGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.