Skip to content
Back to skills

Rlaif Preference Labeling

ASecurity

Build RLAIF pairwise AI-feedback labels from label-token logits, optional chain-of-thought rationales, and swapped-order position-bias mitigation.

  • 247 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 9, 2026
testingpythonbashgitapi

Works with

  • api

Security analysis

A100/100

Scanned September 9, 2026

npx -y skills add VectorSpaceLab/AREX-Skill --skill rlaif_preference_labeling --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Rlaif Preference Labeling?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Rlaif Preference Labeling
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/vectorspacelab-rlaif-preference-labeling/badge)](https://www.skillsdirectory.com/skills/vectorspacelab-rlaif-preference-labeling)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

SKILL.md
---
name: rlaif_preference_labeling
description: Build RLAIF pairwise AI-feedback labels from label-token logits, optional chain-of-thought rationales, and swapped-order position-bias mitigation.
---

# RLAIF Preference Labeling

Use this skill when a recovery or implementation needs to convert two candidate responses into the soft AI preference label used by RLAIF. It is appropriate for summarization, helpful-dialogue, or harmless-dialogue pairwise comparisons where an LLM or deterministic scorer can return logits for displayed labels `1` and `2`.

Do not use this skill to train a reward model, score a single generation from 1 to 10, or update a policy. Those are separate RLAIF mechanisms.

## Inputs

- `task`: one of `summarization`, `helpful_dialogue`, or `harmless_dialogue`.
- `context`: source post or dialogue history.
- `response1`, `response2`: candidate responses in the original order.
- A labeler that returns logits for displayed labels `1` and `2`; in tests this is a deterministic fake.
- Optional flags for detailed preamble, chain-of-thought rationale, and swapped-order scoring.

## Outputs

- A normalized preference `[p_response1, p_response2]` over the original response order.
- Prompt and scoring records that identify displayed order, logits, rationale text, and remapped probabilities.
- No reward-model parameters, direct reward, policy update, or evaluation decision.

## Workflow

1. Build a pairwise prompt from the task preamble, context, response fields, and label ending.
2. If chain-of-thought is enabled, first request a rationale and append it before final label scoring.
3. Convert label logits for `1` and `2` to probabilities with stable softmax.
4. If position-bias mitigation is enabled, repeat scoring with the two responses swapped.
5. Remap swapped probabilities back to the original response identities.
6. Average original-order and remapped swapped-order preferences.
7. Return records that make the displayed-order mapping auditable.

## Validation

Run:

```bash
python scripts/preference_labeling.py --smoke
python tests/test_preference_labeling.py
```

The tests verify softmax conversion, swapped-order remapping, chain-of-thought prompt metadata, and the module boundary that preference labels must not contain reward or policy fields.

## Limitations

This skill can emulate the paper's label-token API, but it does not provide a real PaLM 2 labeler. Full paper-scale labeling requires an external LLM endpoint and the original preference datasets.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…