Skip to content
Back to skills

Assess Construct Validity

ASecurity

Assess whether a benchmark measures its claimed capability using content, convergent, discriminant, and confound analysis.

  • 499 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 24, 2026
research

Security analysis

A100/100

Scanned September 24, 2026

npx -y skills add yogsoth-ai/de-anthropocentric-research-engine --skill assess-construct-validity --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Assess Construct Validity?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Assess Construct Validity
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/yogsoth-ai-assess-construct-validity/badge)](https://www.skillsdirectory.com/skills/yogsoth-ai-assess-construct-validity)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: assess-construct-validity
description: "Assess whether a benchmark measures its claimed capability using content, convergent, discriminant, and confound analysis."
---

# assess-construct-validity

## Purpose

Assess whether a benchmark or evaluation construct measures the claimed capability rather than content cues, confounds, or unrelated skill.

## Input contract

```yaml
required: [construct_claim, benchmark_specification, evaluation_records]
optional: [content_analysis, convergent_measures, discriminant_measures, confound_hypotheses]
constraints: [each validity judgment requires an observable indicator, comparison basis, and provenance]
```

## Procedure

1. State the target construct and map benchmark tasks, labels, and metrics to its intended components.
2. Check content coverage and plausible construct-irrelevant cues against the benchmark specification.
3. Compare convergent and discriminant evidence where available, preserving missing comparisons.
4. Test confound hypotheses with controlled contrasts or artifact probes and record residual uncertainty.

If construct validity depends on whether the operationalization covers the intended domain, consider `map-coverage-space` as the next tactic.

## Output contract

```yaml
produces: [construct_map, content_validity_assessment, convergent_discriminant_evidence, confound_report, validity_judgment]
delta_fields: [findings, evidence_updates, uncertainties, decisions, open_questions]
```

## Quality gates

- Every validity claim cites a task, measure, contrast, or artifact observation.
- Content coverage, convergence, discrimination, and confounds are reported separately.
- A missing diagnostic is marked unresolved rather than treated as evidence of validity.

## Failure and counterexamples

Do not infer construct validity from a high score, face validity, or agreement with another measure that shares the same confound.

## Provenance map

- `resolved: construct-validity-assessment`

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…