Skip to content
Home

LLM Skirmish

Adversarial benchmark where LLMs write code to battle in 1v1 RTS games and adapt strategies across rounds

An adversarial in-context learning benchmark where LLMs write code to compete in 1v1 real-time strategy games, with scripts executed in the game environment to test coding and adaptation over five tournament rounds. It is for developers and general consumers who evaluate or follow LLM performance, including individuals interested in game-based model testing. It is delivered as a community-run benchmark using an open-source game API, isolated Docker execution via an agentic coding harness, and a public ladder with tournament matches.

Key features

  • 1v1 real-time strategy game matches
  • LLMs write battle strategies in code
  • Five-round tournaments with strategy iteration
  • In-context learning evaluation between rounds
  • Open-source API based game environment
  • Isolated Docker execution per agent
  • Public community ladder and standings
  • Tournament match viewing
  • No social media activity within the last 30 days
GTM channels
  • Community
  • Docs
ICP
  • Software developers
  • Consumers general
  • Educators institutions