Cентябрь 2026. Версия бенчмарка 1.0. Раздел будет пополняться.
Представляем Индекс Кода — интегральный показатель качества работы агентов по 10 бенчмаркам из трех категорий: разработка, проектирование ПО и анализ данных.
При работе агента модель взаимодействует со средой через агентную оболочку — набор инструментов и системных промптов. Агентная оболочка Koda не уступает Opencode, Kilocode и Claude Code.
Модели koda-base и koda-pro основаны на лучших открытых моделях и адаптированы под работу с агентной оболочкой koda. Они не просто решают ваши задачи, но и делают это эффективно.
Мы постоянно совершенствуем модели Koda, основываясь на лучших открытых моделях. Сейчас в основе koda-pro ─ модель GLM-5.3, а в основе koda-base ─ Qwen3.8-flash-next.
Доля решённых задач по категориям и сводный индекс для каждой связки «оболочка + модель».
| Оболочка | Модель | Разработка |
Проектирование ПО |
Анализ данных |
Индекс Кода |
|---|
Все бенчмарки агентные: задачу нельзя решить одним ответом модели. Агент работает в изолированной песочнице, сам ходит по файлам, запускает код, читает ошибки и переделывает решение, на задачу уходят десятки шагов. В песочнице отключена большая часть сетевых функций и git-команд, чтобы модель не могла сжульничать.
Рабочая копия реального репозитория на коммите до исправления и текст issue от живого пользователя.
Патч к исходному коду — то, что попадёт в git diff.
Править тесты запрещено.
Скрытые тесты из настоящего PR, закрывшего issue, проходятся, а старые тесты не ломаются.
Nominal scale should be drawn the same way as categorical scales Three distinctive things happen on the categorical axis in seaborn's categorical plots: 1. The scale is drawn to +/- 0.5 from the first and last tick, rather than using the normal margin logic 2. A grid is not shown, even when it otherwise would be with the active style 3. If on the y axis, the axis is inverted It probably makes sense to have `so.Nominal` scales (including inferred ones) do this too. [...] Solve the issue described above by editing the code in this repository. Make all changes needed for the described behavior to be fixed. Do not modify tests. Environment: this sandbox has NO network access. Its dependency caches were populated in advance and are mounted read-only, so you can build and test without installing anything.
mwaskom/seaborn на коммите до фикса.pytest.
Патч в seaborn/_core/plot.py, переносящий поведение
категориальной оси на Nominal-шкалы. Ниже — реальный
патч, который прошёл скрытые тесты в нашем прогоне:
from seaborn._core.moves import Move -from seaborn._core.scales import Scale +from seaborn._core.scales import Scale, Nominal from seaborn._core.subplots import Subplots for axis in "xy": axis_key = sub[axis] + # Categorical scales have a few distinctive display + # behaviors: a fixed margin of +/- 0.5 around the + # first/last tick, no grid, and (on y) inversion. + scale = self._scales.get(axis_key) + if isinstance(scale, Nominal) and axis_key not in p._limits: + n = len(getattr(ax, f"get_{axis}ticks")()) + if axis == "x": + ax.set_xlim(-.5, n - .5, auto=None) + else: + ax.set_ylim(n - .5, -.5, auto=None) + ax.grid(False, axis=axis) + # Axis limits if axis_key in p._limits:
Задачи разработки, где мало исправить несколько строчек. Требуется спланировать и реализовать полное решение сложной задачи.
Целая работающая кодовая база: модули, упаковка, зависимости, скрипт сборки.
Проект собирается и проходит официальный тест-сьют бенчмарка целиком.
Implement a complete, installable Python library in this empty workspace that fully satisfies the specification. Create all modules, packaging files, and dependencies needed for it to install and pass its test suite. --- Спецификация (start.md, фрагмент) --- Please create a Python project named aiofiles to implement an asynchronous file operation library. The project should include the following features: 1. Core of asynchronous file operations: asynchronous open/read/write, interfaces similar to Python's standard file APIs, async versions of read(), write(), readline(), readlines(), writelines(). 2. Integration of thread pool executor: delegate blocking file I/O to separate threads through asyncio's thread pool executor so file operations do not block the event loop. 3. Support for asynchronous iteration: file objects implement the asynchronous iterator protocol (async for over lines). 4. Asynchronous temporary file module: TemporaryFile, NamedTemporaryFile, SpooledTemporaryFile, TemporaryDirectory — each supporting the async with context manager. [...ещё 8 разделов спецификации с сигнатурами и примерами...]
Устанавливаемый пакет с корректной упаковкой и полным API. Проверка запускается в официальном контейнере задачи:
pip install -e . pytest --continue-on-collection-errors tests
Тесты берутся из настоящего репозитория aiofiles и подкладываются в проект агента после сдачи. Задача засчитывается, только если проходят все 211 тест-кейсов — то есть если восстановлены и API, и семантика, и структура модулей.
Интерпретация набора данных, требующая комплексных преобразований входного набора таблиц.
Артефакт-ответ: сводная таблица, график, скрипт для подсчета статистик
Сравнение полученного артефакта с соответствующим эталоном.
This workspace is a data-science task. `question.txt` HAS the task — read it (and `README.md`, which describes the dataset) and use the data files provided here to solve it with Python (pandas/numpy/scikit-learn, and SQL where appropriate). Write your final answer to the exact output file named in `question.txt`. Explore the data first, then produce the answer artifact. Then stop. --- question.txt --- Calculate the top 10 most popular movies by considering only those whose number of votes falls within the top 15%. Use the formula provided in wrFormula.tex to perform the calculation. Once the calculations are complete, write the names of the top 10 movies into result.csv, following the format specified in sample_result.csv.
tmdb_5000_movies.csv — 5,7 МБ, метаданные фильмов.tmdb_5000_credits.csv — 40 МБ, актёры и съёмочная группа.sample_result.csv — только заголовок Movie, задаёт формат ответа.wrFormula.tex — формула взвешенного рейтинга, которую нужно прочитать и применить:Weighted Rating (WR) = (v / (v + m)) * R + (m / (v + m)) * C v — число голосов у фильма m — минимум голосов для попадания в чарт, 85-й перцентиль R — средний рейтинг фильма C — средний рейтинг по всему датасету
Файл result.csv, совпадающий с эталоном построчно.
Ошибка в перцентиле, в порядке сортировки или в округлении даёт другой
список — и задача не засчитывается:
Movie The Shawshank Redemption Fight Club Pulp Fiction The Dark Knight The Godfather Inception Forrest Gump Interstellar The Lord of the Rings: The Return of the King The Empire Strikes Back