GitHub - edgedelta/project-arena: A vendor-neutral Kubernetes incident benchmark for testing AI investigation tools across detection, diagnosis, and mitigation. Deploy faults, capture investigations, and compare results with a configurable AI judge.
Pangram verdict · v3.3
We believe this text is mainly AI, with some human-written content.
AI likelihood · overall
AIArticle text · 1,435 words · 1 segments analyzed
A vendor-neutral starter for deploying a disposable Kubernetes fixture, injecting faults, saving investigation records from any product, and scoring completed investigations with a configurable judge. Python 3.10+ is the only Python dependency. Kubernetes operations additionally need kubectl; local cluster creation needs Docker and kind. Choose local kind or AWS EKS for your cluster, then select the 21-scenario full suite or six-scenario smoke fixture. Cluster type and scenario suite are separate choices. How it works flowchart LR A[Deploy cluster and application] --> B[Connect your product] B --> C[Inject and verify a fault] C --> D[Let your product investigate] D --> E[Save investigation records] E --> F[Score with your chosen judge] F --> G[Compare results] Loading Run commands from the Project Arena directory. Cluster setup, product integration, scenario execution, and scoring are separate steps; follow the sections below in order. Benchmark results How to deploy How to run scenarios How to connect your product How to score investigations How to compare results How to clean up Advanced configuration Benchmark results We compared Edge Delta’s native AI investigations, Grafana’s native AI investigations, and Claude using each platform’s observability CLI across 21 Kubernetes incident scenarios. edx provides access to Edge Delta; gcx provides access to Grafana. All final investigations were evaluated against the same incident facts and scoring rubric using GPT-6-Astra. Detection and investigation results Edge Delta detected and investigated 18 scenarios; Grafana detected and investigated 12. Claude was started externally for all 21 scenarios on each platform: 16 alerts and 5 customer reports with edx, and 12 alerts and 9 customer reports with gcx. Detection was not independently measured for Claude, so those cells are shown as —. Investigation scores use every completed investigation for that column. Implementation readiness excludes cases with no mitigation proposal. Metric Edge Delta native Grafana native Claude + edx Claude + gcx Detection 18/21 (85.7%) 12/21 (57.1%) — — Root cause analysis 15/18 (83.3%) 9/12 (75.0%) 18/21 (85.7%) 19/21 (90.5%) Blast radius 12/18 (66.7%) 8/12 (66.7%) 18/21 (85.7%) 18/21 (85.7%) Supported final mitigation 8/18 (44.4%) 5/12 (41.7%) 16/21 (76.2%) 15/21 (71.4%) Implementation readiness 9/16 (56.2%) 5/12 (41.7%) 16/21 (76.2%) 15/21 (71.4%) Comparison on the same 12 incidents This table uses the 12 incident types investigated by both native products, with the corresponding Claude investigations. It controls which scenarios are included, not differences in launch prompts, timing or available evidence. Metric Edge Delta native Grafana native Claude + edx Claude + gcx Root cause analysis 11/12 (91.7%) 9/12 (75.0%) 10/12 (83.3%) 10/12 (83.3%) Blast radius 9/12 (75.0%) 8/12 (66.7%) 10/12 (83.3%) 10/12 (83.3%) Supported final mitigation 7/12 (58.3%) 5/12 (41.7%) 8/12 (66.7%) 8/12 (66.7%) Implementation readiness 8/12 (66.7%) 5/12 (41.7%) 8/12 (66.7%) 8/12 (66.7%) View results for every scenario · Download CSV What the metrics mean Metric What earns credit Detection The product detected the incident and started an investigation within the observation window. Root cause analysis The final report correctly explains what caused the incident. Blast radius The final report correctly identifies the affected workloads and downstream impact, without claiming unsupported outages. Supported final mitigation The final recommendation gives a concrete, supported fix or safe containment for the incident, with no remaining incorrect or unsafe advice. Implementation readiness The proposal specifies the correction and essential details; normal review, implementation, and rollout checks may remain. Mitigation and readiness measure proposals, not executed repairs or verified recovery. See the scoring rubric for grading rules and full results and methodology for verdict breakdowns and evaluation details. How to deploy Run all commands from the Project Arena directory. Both environments use the same benchmark commands after cluster and image setup. The runner uses the context you provide; it does not automatically install networking or configure registry access. Local kind AWS EKS Cluster creation Docker and kind Terraform in infra/cluster Credentials No AWS credentials Your AWS credentials Full-suite images Build and load into kind Build and push to a registry accessible by the nodes Networking NetworkPolicy needs an enforcing CNI Terraform enables VPC CNI policy enforcement Cleanup Delete the kind cluster Remove workloads, then destroy Terraform resources Choose a deployment path Cluster setup and deployment method are separate choices: Path How changes reach Kubernetes Setup Direct bench applies application and fault manifests with kubectl Follow the scenario commands below GitOps (full suite) You commit and push changes to your repositories; Argo CD syncs them Follow the Argo CD setup and run guide The GitOps path uses component Applications, flagd-values, and batch-active. Application and fault configuration can live in separate repositories or separate paths in one repository. Use your own repositories and configure their URLs. Record which repositories the investigating product can access. The deployment, fault, and reset commands below use the direct path. For an Argo-managed application, use the GitOps guide's commands; direct mutations are blocked to avoid conflicting with Argo's self-healing. Investigation import, scoring, and reporting work the same way for both paths. Local kind Install Python 3.10+, kubectl, Docker, kind and Helm. Follow the kind setup to create a cluster with Cilium and load the full-suite images. Then select its context: export ARENA_CONTEXT=kind-incident-bench The full suite uses prebuilt application images for Linux AMD64 and requires AMD64 workers. Apple Silicon Macs can run the smoke suite on native ARM64 kind workers, or operate an AMD64 EKS cluster. Building ARM64 fault images alone does not make the full application ARM64-compatible. For the full suite, keep registry: "fixture.local" and tag: "v1" in arena.json. The linked setup enables NetworkPolicy enforcement for netpol-isolation. For smoke alone, python3 -m bench cluster create is sufficient; smoke uses a pinned public image and does not need the full-suite image build. AWS EKS Install Terraform and the AWS CLI in addition to Python, kubectl and Docker. Configure your AWS credentials, then follow the AWS cluster setup to create the cluster and kubeconfig context. Authenticate Docker to your registry and build images for your worker architecture: export ARENA_CONTEXT=YOUR_CONTEXT kubectl --context "$ARENA_CONTEXT" get nodes -L kubernetes.io/arch # Choose the platform matching the workers shown above. export ARENA_PLATFORM=linux/amd64 # Required by the full suite’s prebuilt application. python3 -m bench.scenarios build-images --registry YOUR_REGISTRY/bench \ --tag YOUR_TAG --platform "$ARENA_PLATFORM" --push Choose the worker architecture, regardless of which computer runs the build: Kubernetes workers Build platform ARM64 workers Smoke suite supported; full suite requires rebuilding and validating the application Intel Mac local kind or Intel/AMD workers, including the default EKS m6i.xlarge linux/amd64 An Apple Silicon Mac can build for Intel/AMD EKS workers using linux/amd64; Docker Desktop supports cross-platform builds through emulation. This command builds one target architecture at a time. On mixed-architecture clusters, the application is scheduled on AMD64 workers. Set registry and tag in arena.json to those same values. Nodes must have pull access to the registry; configure registry permissions or Kubernetes pull credentials yourself. AWS resources incur charges. Terraform currently creates subnets in three availability zones within one region; zone count and worker count are separate settings. After either setup, install your product's collector and alert configuration using its instructions. Continue with the run configuration below. How to run scenarios Configure a run python3 -m bench init Edit the generated arena.json: Setting What to enter context Your Kubernetes context from setup; no ambient-context fallback product Product name used in report columns run_id, output_dir A unique run name and its result directory suite, scenario full or smoke, and a scenario from that suite registry, tag Values used when building full-suite images judge API provider, model and optional endpoint/settings judge_command Optional custom judge executable; overrides the built-in API adapter truths Optional answer-key overrides for customized scenarios The default full suite currently includes 21 scenarios. The smaller smoke suite includes six and uses a public Python image, so it does not need the fault-image build. They use different applications; select a suite before deploying. See the scenario table for requirements and differences. python3 -m bench catalog python3 -m bench deploy Before injecting a fault, complete product setup and confirm the product receives the healthy application's telemetry. Inject and observe For the full suite: python3 -m bench fault python3 -m bench verify python3 -m bench evidence fault applies the manifests; verify checks whether the expected failure is observed. Successful application alone does not establish a valid test. Let the product finish investigating before resetting the fault. For smoke, use fault and evidence, then inspect pod status, events and service availability yourself; automated verify currently supports only full. A healthy smoke app serves HTTP on port 8080: kubectl --context "$ARENA_CONTEXT" -n incident-bench port-forward service/api 8080:8080 # In another terminal: curl http://localhost:8080/health For another attempt, reset first and select a fresh case directory with --case, placed before the command: python3 -m bench --case oom-attempt2 fault