We’re an independent team of cybersecurity experts building original, verified datasets for companies developing local or specialized models, frontier labs, and data-labeling providers. Every task comes with a reproducible environment, a verified solution path, and reward checks you can trust.
Drag upward to separate the layers, or use the up and down arrow keys. Drag sideways or use the left and right arrow keys to rotate. Press Home to reset. The production stages advance automatically.
We tailor the domains, difficulty, environment, output format, and validation depth to your project. If you already handle annotation and ML preparation, we provide the security layer: original tasks, verified solutions, reviewed agent runs, and reliable reward logic.
Every set includes original tasks, verified solutions, reviewed agent runs, and reward checks that hold up.
Reviewed by security experts. Delivered privately.
Built around your requirements.
Choose the parts you need or ask us to handle the full dataset. We tailor the task package, solution evidence, agent traces, reward checks, and review notes to your workflow.
01
Task and environment
The objective, constraints, and everything needed to reproduce the task.
02
Verified solution
A tested solution path with the reasoning and evidence behind it.
03
Agent run review
What agents tried, where they got stuck, and which failures matter.
04
SFT preparation
Clean reference trajectories that can be used for supervised fine-tuning.
05
Reward checks
Tests for real task completion, shortcuts, and evaluator exploits.
06
Expert review
A second opinion on difficulty, fairness, and training value.
Hard tasks. Rewards that hold up.
We reproduce the solution, run agents through the task, and try to break the reward logic. If the task tests the wrong thing, we send it back.
Illustrative review
The task works as intended.
The agent follows a reproducible path and completes the actual objective. A reviewer confirms the task is fair.
Solution path
A solution another expert can reproduce.
Reproducing the solution…
Agent run
Where the agent gets stuck and why.
Awaiting review
Reward integrity
Reward the objective, not a shortcut.
Awaiting review
Expert review
Hard for the right reasons.
Awaiting review
Review in progress
Checking the solution, agent run, and reward logic.
The task works as intended.
The agent follows a reproducible path and completes the actual objective. A reviewer confirms the task is fair.
Solution path: Solution reproduced.
Agent run: Agent run reviewed.
Reward integrity: Real objective completed.
Expert review: Difficulty looks fair.
Ready for delivery. The solution works, the agent run has been reviewed, and the reward matches the objective. It’s ready for training or evaluation.
THE TEAM BEHIND THE TASKS
Security experts. Dataset builders.
Our founders combine hands-on security research with experience in data labeling, dataset review, CTF operations, reward design, and LLM research.
0xp1ain leads RewriteLab, an international web security research team. He has worked on cybersecurity data labeling for multiple providers, benchmark dataset quality review, high-difficulty CTF design and operations, and LLM research. At CYVARC, 0xp1ain leads fair task design and end-to-end validation.
sahuang is the lead AI security researcher at OtterSec, founder of the internationally recognized CTF team Project Sekai, and a former Microsoft engineer. Drawing on experience leading cybersecurity dataset production and competing in international CTFs, sahuang leads dataset structure, reward integrity, and technical review at CYVARC.
For companies developing local models, frontier labs, and specialist data-labeling providers.
FOR COMPANIES BUILDING LOCAL MODELS
Train your model on tasks built for its needs.
We tailor the domains, difficulty, tool access, context, and task volume to your training setup, then deliver the data in the schema your pipeline expects.
FOR FRONTIER LABS
Test the edge of capability.
Commission private, high-difficulty tasks with reproducible environments, verified solution paths, reviewed agent runs, and reward logic tested for shortcuts.
FOR DATA-LABELING COMPANIES
Add the cybersecurity layer.
We plug into existing labeling and post-training workflows for task writing, domain review, solution validation, reward design, and quality control.
Made to your specification.
Set the domains, difficulty, volume, environment, output schema, and level of review. We can deliver task packages, reviewed agent traces, SFT-ready trajectories, reward checks, or a complete dataset.
Common questions.
CYVARC is an independent team, not a registered legal entity. We’re cybersecurity researchers who come together to build original tasks for AI training and evaluation. Each project is scoped separately, and we bring in the people it needs.
A good CTF challenge is only the starting point. We also verify the intended solution, run agents against the task, look for broken assumptions, test the reward logic, and document what we find. That turns a challenge into training data companies and labs can actually use.
We work across binary exploitation, reverse engineering, cryptography, cloud security, smart contract security, and web security. For each project, we define the target skill first, then bring in people who know that area well.
We start with a simple question: does the task teach the skill it claims to test? Our reviewers reproduce the solution, inspect agent runs, check for unclear instructions or fragile environments, and make sure the reward can’t be earned through a shortcut.
We work with companies developing local or specialized models, frontier labs, and data-labeling or post-training providers that need cybersecurity expertise for their training or evaluation data. We can build the full dataset or fit into an existing annotation and ML pipeline.
Yes. We can match your target capabilities, difficulty range, task volume, agent setup, output schema, and review requirements. Deliveries are private, and you can request anything from standalone task packages to SFT-ready traces and reward checks.
Need verified cyber tasks? Let’s scope your dataset together.