Skip to content
Back to skills

Research Engineer

ASecurity

Do applied AI research engineering — reproduce papers, run ablations, manage experiments, and turn prototypes into reliable systems. Use when bridging the gap between a research idea and working code.

  • 2 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 29, 2026
ai-agentsgodebugging

Security analysis

A100/100

Scanned September 29, 2026

npx -y skills add aicodedecode/awesome-muse-skills --skill research-engineer --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Research Engineer?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Research Engineer
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aicodedecode-research-engineer/badge)](https://www.skillsdirectory.com/skills/aicodedecode-research-engineer)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: research-engineer
description: Do applied AI research engineering — reproduce papers, run ablations, manage experiments, and turn prototypes into reliable systems. Use when bridging the gap between a research idea and working code.
category: ai-research
---

# Research Engineer

The research engineer lives between the paper and the product: reproducing results, running clean 
experiments, ablating ideas, and hardening prototypes. The craft is rigor at speed — scientific 
discipline without academic timescales.

## Overview

A research idea is a hypothesis until it's reproduced and ablated. The workflow: understand the 
paper deeply enough to reimplement from scratch, reproduce the key result, then vary one thing at a 
time to learn what actually matters. Experiment management — configs, seeds, logs, versioning — 
is what separates real findings from anecdotes. The output is knowledge (what works, what doesn't, 
why) plus code others can build on.

## When to use

- A paper claims a result you want to use or build on — reproduce it first.
- Comparing approaches: which component actually drives the gain?
- Turning a promising prototype into something reliable enough to ship.
- Debugging why a published method doesn't work in your setting.

## Core concepts

- **Reproduction**: reimplementing from the paper (not just running released code) to truly 
understand it. If you can't reproduce it, you don't understand it.
- **Ablations**: removing or varying one component at a time to measure its contribution. The 
single most informative experiment type.
- **Experiment tracking**: every run logged with config, code version, seed, data version, and 
metrics. Untracked experiments are rumors.
- **Baselines**: the simplest alternative that could work. New ideas must beat strong baselines, 
not strawmen.
- **Statistical honesty**: multiple seeds, error bars, and pre-registered comparisons. One lucky 
seed is not a result.
- **Negative results**: documenting what didn't work is as valuable as what did. Keep a lab 
notebook of failures.

## Practical workflow

1. Read the paper for the mechanism, not the marketing: what is the actual claim, and what would 
falsify it?
2. Reimplement the core idea minimally; reproduce the headline result on the paper's setup before 
adapting it.
3. Establish baselines on your data/task — including the "dumb" approaches.
4. Ablate: vary one factor at a time (architecture, data, hyperparameters, prompts) and record the 
delta.
5. Stress-test the winner: distribution shift, edge cases, ablations of your own additions.
6. Write it up: method, setup, results with uncertainty, what failed, and what you'd try next.

```text
Experiment log template:
HYPOTHESIS: <what you expect and why>
SETUP:      <config, code commit, seed, data version>
BASELINE:   <score to beat>
RESULT:     <score ± spread over seeds>
DELTA:      <vs baseline and vs ablations>
VERDICT:    <keep / kill / investigate>
```

## Common pitfalls

- **Running released code only**: you learn to operate it, not understand it. Reimplement the core.
- **Weak baselines**: beating a poorly tuned baseline proves nothing. Tune baselines as hard as 
your idea.
- **Single-seed results**: reporting one lucky run. Use multiple seeds and report spread.
- **Ablation theater**: ablating trivial variants while the real question goes untested. Ablate the 
claim, not the decoration.
- **Untracked experiments**: "I think we tried that." Log everything; memory lies.
- **HARKing**: hypothesizing after results are known. Write the hypothesis before running, and 
honor negative results.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…