CybORG
Simulated cyber-defense arena based on the CAGE Challenge 3 DroneSwarm scenario.
Overview
CybORG is a cyber operations research gym for training and evaluating autonomous security agents. The CodeClash arena uses CybORG's simulated DroneSwarm scenario through the PettingZoo parallel interface. It does not run real exploit tooling, emulate external networks, or interact with live systems.
Each CodeClash player edits a restricted blue-team policy. A round evaluates every submitted policy on the same seeded episode batch and scores players by average episode reward. The trusted runtime owns the CybORG environment and action validation; submitted code only receives plain observations and returns discrete action intents.
Resources
Implementation
codeclash.arenas.cyborg.cyborg.CybORGArena
CybORGArena(config: dict, *, tournament_id: str, local_output_dir: Path, keep_containers: bool = False)
Bases: CodeArena
Source code in codeclash/arenas/arena.py
83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 | |
name
class-attribute
instance-attribute
name: str = 'CybORG'
submission
class-attribute
instance-attribute
submission: str = 'cyborg_agent.py'
description
class-attribute
instance-attribute
description: str = 'CybORG is a simulated cyber-defense arena based on the CAGE Challenge 3 DroneSwarm scenario.\n\nYour bot is a Python file named `cyborg_agent.py` that defines a function named `decide`.\nThe function receives a plain observation list and action-space dictionary, then returns an action:\n\n def decide(observation, action_space):\n return 0\n\nEach round evaluates every submitted agent independently on the same seeded DroneSwarm episodes.\nThe trusted runtime owns the CybORG environment and action validation. Submitted code only receives\nplain observations and returns action intents. The objective is to maximize average episode reward.\nThis arena uses CybORG simulation only and does not run real exploit tools or interact with external\nnetworks.\n'
default_args
class-attribute
instance-attribute
default_args: dict = {'steps_per_episode': 30, 'num_drones': 18, 'decision_timeout': 3.0, 'validation_timeout': 10, 'timeout': 240}
validate_code
validate_code(agent: Player) -> tuple[bool, str | None]
Source code in codeclash/arenas/cyborg/cyborg.py
45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 | |
execute_round
execute_round(agents: list[Player]) -> None
Source code in codeclash/arenas/cyborg/cyborg.py
81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 | |
get_results
get_results(agents: list[Player], round_num: int, stats: RoundStats)
Source code in codeclash/arenas/cyborg/cyborg.py
109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 | |
Agent Interface
Your bot must be a Python file named cyborg_agent.py that defines decide(observation, action_space).
observation is a plain list converted from CybORG's observation array. action_space is a
dictionary such as {"type": "discrete", "n": 11}. Return an integer action accepted by that
action space. A valid starting point is:
def decide(observation, action_space):
return 0
For each episode, the same policy file controls all blue-team drone agents through isolated policy worker processes. The runtime validates every returned action before stepping the simulator.
Configuration Example
tournament:
rounds: 1
game:
name: CybORG
sims_per_round: 2
args:
steps_per_episode: 5
num_drones: 8
decision_timeout: 3.0
validation_timeout: 10
timeout: 240
players:
- agent: dummy
name: alpha
- agent: dummy
name: beta
Scoring
The arena runs sims_per_round independent simulated DroneSwarm episodes for each submitted player.
Each player receives the sum of mean blue-agent rewards per episode. The final CodeClash score is the
average episode score across the round.
The runtime pins CybORG to the upstream v3.0 code and installs it editable from a checked-out
repository because the upstream package expects data files such as CybORG/version.txt to be present
next to the source tree.
Smoke Test
From the repository root, run the dummy-player example:
uv run codeclash run configs/examples/CybORG__dummy__r1__s2.yaml -o /tmp/codeclash-cyborg-smoke
Use a fresh -o directory when rerunning the smoke check.
Expected shape:
- the command exits with status 0;
- both players pass submission validation;
- stdout includes
In round 0, the winner is ...andIn round 1, the winner is ...; - each round summary contains floating-point average rewards for
alphaandbeta; - per-episode details have
status: "ok",steps_completed: 5,policy_errors,invalid_actions, anddecisions; - the output directory contains
metadata.json,game.log,tournament.log, androunds/round_0.tar.gz/rounds/round_1.tar.gz.
A representative metadata.json round contains a scores object with one floating-point episode
reward per player:
"scores": {
"alpha": -27.0,
"beta": -33.5
}
Exact values can change with simulator randomness and configuration; the smoke check is meant to verify the Docker/runtime adapter path, player-name mapping, and score/log artifact shape.
The exact tournament directory name includes a timestamp, so inspect the metadata with:
find /tmp/codeclash-cyborg-smoke -maxdepth 3 -name metadata.json -print