LLMs exhibit intelligence, they have solved open Maths problems, formalized Fermat's Last Theorem and have consumed more internet than any human alive. But day-to-day, why can't we get them to do things the way we like, no matter how many skills we write?

To understand this, I propose a though experiment. Let's say from today onwards, every morning, your body will reset back to 6th of September 2026. Anything you have learned after that point, any skills, facts or muscle memories will get wiped out as soon as you go to bed.

Rest of the world functions as-is. If you make a financial investment, it'll still be there. If write something in the Diary, you can read it the next day. But your body and mind will always start as a clone of 6th-of-September-2026-you.

Given this setup, can you learn to drive (assuming you don't already know it)? After a driving lesson, you can try to take notes so you don't have to start from scratch the next day. But will those help? Let's say your notes look like:

  • Check rear mirror before turning.
  • I was 300ms late in changing the gear, I should be faster.
  • I was too fast when driving downhill.

We know, that we cannot learn driving if all we have is these notes. We learn driving by accumulating experiences, not with data and instructions.

This is the life of an Agent. It has some context limit, after which it has to restart its day with just the "notes".

This is why I believe Agent Skills is a misleading term. Agents do not genuinely aquire skills. A better way to think about these "skills" is to just think of them as notes.

Once we understand this, it becomes easy to understand why a lot of pattern that feel like they should work, start to fail in practice. For example: "Hey agent, whenever you learn something about my preference, write it to memory.md". With experience, I know that over time memory.md results in a very rigid set of instructions which reflect my preferences purely. This is not say, this pattern isn't helpful at all.

In driving context, memory.md is useful for:

  • The road remains closed on sunday
  • Going via the road near park saves the time due to less traffic

But it is not helpful for the muscle memory. So agent will be able to learn which tools it should use, but it will likely struggle with learning the nuances of a particular instruction.

For example, let's say you ask an Agent to review a piece of text. It comes out by pointing 5 grammatical mistakes. Then you point out that it should focus more on the content and less on the grammatical mistakes. The agent might take a note that "User prefers the review to focus more on the content and less on the grammar".

Later, you ask it to review another text and it misses an obvious mistake. You point out that certain sentences aren't structurally valid. It might take a note of "User prefers the review to focus more on the content and less on the grammar unless the sentences aren't structurally valid".

A few examples of what is likely to work if you put it in SKILLS.md:

  • Always summarize output in 1 sentence
  • Use human readable numbers like 10k instead of 10000

A few examples that will likely not provide intended effect:

  • Provide wikipedia links for vague terms: It will fail because LLM cannot know what is vague for you.
  • Always review the changes holistically: It will fail because holistically is harder to define. Holistically at the scale of company, country, human race or universe?

We know that it still doesn't capture the intent of what we consider to be a good review. A review should cover all the cases, balancing various aspects. But because the agent's memory is restricted by english notes, it cannot accumulate the experiential knowledge.

Using examples is a good way to get around some of these limitations, but in some cases the limitation is fundamental and we might need to reach out to complex orchestrations with multi-agent workflows to achieve what we want.