All authors

Claude Skills by yogsoth-ai
github.com/yogsoth-ai1,223 skills7 installs1,449 views
- Injection FidelityLoss-1 judge (codex role). Given one sample's de-identified dialogue and its PolicyCard, decide axis-by-axis whether the user-simulator enacted the card's per-axis pressure. Judge enactment of the card, never whether the research is good.Votes: 0GitHub stars: 417
- Ladder Quality OrderLoss-2 judge (codex role). Over one topic's 6 shuffled research-design samples, pairwise-rank by quality using the D1–D5 standard. Emit the pairwise log; the harness computes the order and the ladder verdicts. Judge quality difference, never against academic standards.Votes: 0GitHub stars: 417
- Optimization LoopThe optimizer brain for the ladder-foundry pretraining loop. Runs the two-level nested batch loop, delegates gating to gate_eval, attributes a failing batch to one weight (attribute-first), and recovers from disk after compaction. Control flow is fully scripted; only the backprop attribution is a judgment call.Votes: 0GitHub stars: 417
- Acu Nugget RecallTactic: Extract atomic units from one paper and score how much of a caller-supplied summary covers. Use for ACU-style binary or Nugget-style ternary recall checks; cannot run without a target summary.Votes: 0GitHub stars: 417
- Argumentative ZoningTactic: Label every sentence of one paper with its rhetorical role using Argumentative Zoning. Use when fixed rhetorical labels and cross-paper alignment matter.Votes: 0GitHub stars: 417
- Atomic Unit MatchingJudge, per atomic content unit, whether a target text (summary, abstract, or other candidate text) contains it — binary present/absent (ACU) or ternary support/partial_support/not_support (Nugget), per caller's value domain. Use this after atomic-unit-writing has produced the reference units, as the matching step before recall aggregation.Votes: 0GitHub stars: 417
- Atomic Unit Recall AggregateAggregate per-unit ACU/Nugget match judgments into a final recall score — normalized length-penalized recall for ACU, or V_strict/A_strict (+ run-level ranking, with an explicit per-topic-unreliability caveat) for Nugget. Use this as the final step of the atomic-unit chain, after atomic-unit-matching; this SOP's existence closes a gap the original pipeline design was missing — without it, per-unit match judgments were never actually summed into the score the source methodologies report.Votes: 0GitHub stars: 417
- Atomic Unit WritingExtract (ACU-style) or freshly author (Nugget-style) a list of atomic content units from a paper, optionally tagged vital/okay for importance. Use this as the first step whenever building a reference set of atomic facts for later recall-checking a summary or abstract against the paper — always precedes atomic-unit-matching.Votes: 0GitHub stars: 417
- Claim Label PredictionJudge a three-way SUPPORTS/REFUTES/NOINFO label for an atomic claim, based only on its selected rationale sentences (SciFact's final classification step). Use this after rationale-selection has produced the evidence sentences — this is the terminal step of the SciFact chain, producing the complete (claim, abstract, label, rationale) tuple.Votes: 0GitHub stars: 417
- Claim WritingBlind-rewrite a citing sentence (citance) from another paper into a single atomic, independently-verifiable claim (SciFact's annotation protocol) — never looking at the cited paper's content while rewriting. Use this when you have a specific citing sentence and want it decomposed into checkable atomic claims, as the first step before rationale-selection and claim-label-prediction.Votes: 0GitHub stars: 417
- Domain Level JudgmentFold raw signalling-question answers into domain-level judgments for RoB2, ROBINS-I, or QUADAS-2, per each tool's own lookup rules — the first of two aggregation levels these tools define. QUADAS-2 is dual-axis (risk-of-bias AND applicability-concern per domain, D1-D3) and terminates here with no further rollup; RoB2/ROBINS-I continue on to worst-case-lookup for an overall verdict. Use this after signalling-question-answering has produced the raw answers.Votes: 0GitHub stars: 417
- Dual Column Self CheckRun one of the ML/CS reproducibility checklists (ML Reproducibility Checklist, REFORMS, NeurIPS Paper Checklist, Model Cards, Datasheets for Datasets) against a paper as a reader-side audit, producing a category (Yes/No/NA) plus free-text reason per item. Use this whenever the user wants a reproducibility/completeness self-check run on an ML or CS paper — invoke this directly, it has no study-design gate in this package since these checklists are engineering self-audits, not clinical-study to...Votes: 0GitHub stars: 417
- Engineering Config Grading(Proposal, unverified) Grade reproducibility-relevant engineering configuration items (hyperparameter search range, compute budget, seed handling, dataset splits) on a complete/partial/none scale, requiring the grader to first define what "complete" means per item before judging against it. Use this after study-design-tool-gate has dispatched an ML/CS engineering paper here; this is a graded QUALITY judgment, distinct from dual-column-self-check's binary Yes/No/NA self-audit checklists.Votes: 0GitHub stars: 417
- First Pass SkimKeshav's first pass over one paper — a 5-10 minute skim of title, abstract, headings, figures, and conclusion only, producing skim notes and a read-deeper judgment. Use this as the first step whenever a paper is being read via the Keshav three-pass method; always precedes second-pass-grasp and never reads section bodies itself.Votes: 0GitHub stars: 417
- Keshav Three PassTactic: Read one paper by Keshav''s three-pass method — a shallow skim, a contribution-grasping full read, then a deep virtual re-implementation. Use when the goal is understanding a paper rather than extracting a fixed schema.Votes: 0GitHub stars: 417
- Multi Stage Cascade ExtractionRun a multi-stage extraction cascade (mention detection, document-level coreference clustering, optional saliency judgment, N-ary relation/triple extraction) directly over a paper's full text — covers SciERC, SciREX, and NLP Contribution Graph. Use this whenever cross-sentence or document-level entity/relation extraction is needed (e.g. SciREX-style Task-Dataset-Metric-Score tuples); do NOT use unit-classification for this, since these methods reason over the whole document's mentions, not in...Votes: 0GitHub stars: 417
- Paper FetchRetrieve one specified academic paper (by title, arXiv ID, DOI, URL, or a local .md/.txt/.pdf path the caller already has) and land it on disk as source.md plus a source.meta.json carrying a line-number section index. Checks context/papers/ for an existing copy first; local files and direct PDF URLs are read directly with no search at all, while other references use alphaxiv, Semantic Scholar routing, then bioRxiv/medRxiv. Use this as the mandatory first step whenever any other paper-reading ...Votes: 0GitHub stars: 417
- Qalmri WorksheetTactic: Fill a six-slot QALMRI worksheet for one paper: Question, Alternatives, Logic, Method, Results, and Inference. Use for a structured reading worksheet rather than a graded evaluation.Votes: 0GitHub stars: 417
- QalmriProduce a six-slot QALMRI worksheet (Question, Alternatives, Logic, Method, Results, Inference) as free-text notes on one paper — a structured note-taking format, not a scored evaluation. Use this whenever the user wants a QALMRI-style reading worksheet for a specific paper.Votes: 0GitHub stars: 417
- Qasper Evidence QaAnswer a specific question about a paper, grounding the answer in exact quoted evidence spans from the text (QASPER-style question-driven QA with span-level evidence, no schema categorization). Use this whenever the user asks a specific factual question about a paper and wants the answer traceable to exact text spans.Votes: 0GitHub stars: 417
- Quality Appraisal ChecklistRun CASP (8 study-type variants), JBI (~6 variants), or AMSTAR-2 quality-appraisal checklists — each ending in the tool''s own required integrated judgment, not just item tallies. Also runs a proposal "rhetorical-completeness-check" mode (entry_mode="completeness_check") that instead diffs unit-classification''s rhetorical labels against a target checklist''s expected label set. Use this after study-design-tool-gate has dispatched to CASP/JBI/AMSTAR-2 (mode a), or directly after unit-classifi...Votes: 0GitHub stars: 417
- Question FramingFill a slot-based question-framing schema (PICO, PECO, or SPIDER) from a paper's stated research question. Use this whenever the user wants a paper's research question structured into one of these standard clinical/qualitative-research question frames; this frames what question is being asked, it does not read or evaluate the paper's content otherwise.Votes: 0GitHub stars: 417
- Rationale SelectionSelect the minimal set of 1-3 verbatim sentences from a candidate paper/abstract sufficient to entail or refute an atomic claim (SciFact's rationale-selection step). Use this after claim-writing has produced an atomic claim, as the evidence-gathering step before claim-label-prediction; an empty rationale set is a valid outcome, not an error.Votes: 0GitHub stars: 417
- Reforms GradingTactic: Grade an ML/CS paper''s reproducibility configuration reporting as complete, partial, or none after checking that clinical appraisal tools do not apply. Use when the question is whether the work can be rerun.Votes: 0GitHub stars: 417
- Reporting Standard ChecklistCheck whether a paper reports each item from PRISMA, CONSORT, STROBE, ARRIVE, SPIRIT, or TRIPOD (per whichever study-design-tool-gate dispatched to), citing where each item is or isn't addressed — including a/b sub-item hierarchy where the standard defines one. Use this after study-design-tool-gate has dispatched to one of these 6 reporting standards; this checks report completeness (did they say where), not methodological quality (was the study done well) — there is no overall synthesis step...Votes: 0GitHub stars: 417
- Reproducibility Third Party Verification(Proposal, unverified) Attempt to verify a paper's reported results by actually executing its released code/scripts against its own reported configuration — the only SOP in this package whose action type is code execution rather than text reading/judgment. Use this after unit-classification has extracted the paper's reported configuration/hyperparameters as classified units; "not_attempted" is a correct, common output when the paper's own reporting is too incomplete to run, not a failure of t...Votes: 0GitHub stars: 417
- Research Question AppraisalJudge a paper's stated research question against the FINER criteria (Feasible, Interesting, Novel, Ethical, Relevant) — five independent judgments with justification, evaluating the question itself, not the paper's results. Use this whenever the user wants to know whether a paper is asking a good research question, distinct from whether it answered that question well.Votes: 0GitHub stars: 417
- Rhetorical Structure Quality(Proposal, unverified) Judge whether argumentative relations between unit-classification's rhetorical labels actually hold in a paper (e.g. is an AIM label adequately substantiated by BACKGROUND labels) — a second-order quality judgment over already-classified units, not raw text. Use this after unit-classification has labeled a paper's units with a rhetorical/argumentative label set, when the user wants to know if the paper's argument structure is actually sound, not just what role each sent...Votes: 0GitHub stars: 417
- Second Pass GraspKeshav's second pass — a careful full read (ignoring proof/derivation detail) producing prose-level understanding sufficient to explain the paper's main contribution and evidence to a colleague. Use this after first-pass-skim, as the main content-grasping pass of the Keshav three-pass method; do not force its output into a structured data schema.Votes: 0GitHub stars: 417
- Signalling Question AnsweringAnswer per-domain signalling questions (5-value scale: Yes/Probably yes/Probably no/No/No information) for RoB2, ROBINS-I, or QUADAS-2, per whichever variant study-design-tool-gate dispatched to. Use this after study-design-tool-gate has dispatched to one of these three tools; this SOP produces only the raw signalling answers, not any domain-level or overall roll-up — that happens in domain-level-judgment next.Votes: 0GitHub stars: 417
- Star AwardingAward NOS's (Newcastle-Ottawa Scale) stars item-by-item across Selection (up to 4), Comparability (up to 2), and Outcome/Exposure (up to 3) — a binary award-or-not action per item, distinct from a 5-value signalling judgment. Use this after study-design-tool-gate has dispatched to NOS, as the first step before sum-threshold-scoring.Votes: 0GitHub stars: 417
- Study Design Tool GateClassify a paper's study design (RCT, cohort, case-control, diagnostic-accuracy, systematic-review, animal-study, prediction-model, etc., or not_applicable) and dispatch to the correct downstream bias-risk/quality/reporting tool and specific variant (CASP has 8 variants, JBI ~6, RoB2 has parallel/cluster/crossover versions). Use this as the mandatory first step before running ANY of CASP, JBI, AMSTAR-2, NOS, RoB2, ROBINS-I, QUADAS-2, CONSORT, STROBE, ARRIVE, SPIRIT, TRIPOD, or engineering-con...Votes: 0GitHub stars: 417
- Sum Threshold ScoringSum NOS's item-level stars and bucket into good (≥7)/fair (4-6)/poor (≤3) — a fixed threshold lookup, structurally distinct from worst-case-lookup's take-the-worst-value approach. Use this after star-awarding has produced the per-item stars; this is NOS's terminal step.Votes: 0GitHub stars: 417
- Template Slot FillingFill a paper's reported values into an already-given comparison-template attribute schema (e.g. Task/Dataset/Metric/Value) — the executable half of ORKG's comparison-template method. Use this when a template's attribute schema is already fixed and you need one paper's row filled in; this does NOT build new templates (that half is a human-curator task, out of scope).Votes: 0GitHub stars: 417
- Third Pass Deep ReadKeshav's third pass — the heaviest of the three, a full sentence-by-sentence re-read including proofs/derivations, attempting a virtual re-implementation of the paper to surface implicit assumptions and concrete improvement points. Use this after second-pass-grasp, as the terminal step of the Keshav three-pass method, whenever genuine mastery of a paper (not just a summary) is needed. This is not a skippable recap — treat "nothing new to add" as suspicious, not a default outcome.Votes: 0GitHub stars: 417
- Unit ClassificationClassify each pre-segmented text unit independently against a fixed label set (Argumentative Zoning, CoreSC, PubMed-RCT, Swales move/step, CODA-19, TDMS, or CSFCube's facet labels), single-layer with no cross-unit dependency. Use this after unit-segmentation has split the text, whenever a sentence- or clause-level rhetorical/functional classification is needed; do not use this for methods requiring document-level coreference reasoning (see multi-stage-cascade-extraction instead).Votes: 0GitHub stars: 417
- Unit SegmentationSplit a paper's text into sentence- or clause-level units (with character offsets) for downstream classification, at a caller-specified granularity and scope (full text, abstract-only, or intro-only). Use this as the mandatory first step whenever any sentence/clause-level classification method (Argumentative Zoning, CoreSC, PubMed-RCT, CSAbstruct, Swales move analysis, CODA-19) needs its input pre-segmented — always precedes unit-classification.Votes: 0GitHub stars: 417
- Worst Case LookupTake the single most severe domain/item judgment as the overall verdict, for RoB2 (3-value), ROBINS-I (5-value), or AMSTAR-2 (pre-filtered by critical-domain status before worst-case). Use this after domain-level-judgment (for RoB2/ROBINS-I) or quality-appraisal-checklist (for AMSTAR-2) has produced per-domain/item judgments — this SOP has two structurally distinct upstream callers and must identify which value domain it received before applying the matching lookup rule. QUADAS-2 never reache...Votes: 0GitHub stars: 417
- Repo Dependency GraphReconstruct a DARE skill repo's true use-dependency relations and render them as a self-contained, offline, Obsidian-style interactive HTML graph (pyvis / vis-network). Use this whenever the user wants to graph / map / visualize the skill dependencies of a repo or package, "画依赖图 / graph 化这个 repo / 把 skill 连边画出来 / 用 pyvis 出个图 / skill 关系图", or to audit how campaign→strategy→ tactic→sop skills connect. Trigger even if the user just says "给这个 package 做个图" without naming pyvis or HTML. Goes straig...Votes: 0GitHub stars: 417
- Adversarial Debate Truthseeking'Strategy: Dialectic engine retuned for truth-seeking, not survival.Votes: 0GitHub stars: 417
- Alias ResolutionSOP for detecting and resolving concept aliases — merge duplicate pages,Votes: 0GitHub stars: 417
- Ara CompileSOP: Turn the feeding plan into the compiler''s $ARGUMENTS and run the external ARA compiler once inline to produce ../ara/Votes: 0GitHub stars: 417
- Ara From ContextCampaign: Compile a context/ research record into an ARA (Agent-Native Research Artifact) and run a Level-2 epistemic review — no LaTeX, no narrative paperVotes: 0GitHub stars: 417
- Ara Rigor ReviewSOP: Run the external ARA rigor-reviewer (Seal Level 2, six-dimension semantic review) over ../ara/ and pass its level2_report.json to the userVotes: 0GitHub stars: 417
- Argument MappingCampaign for mapping argument structures — extract claims, link evidence,Votes: 0GitHub stars: 417
- Argument SynthesisStrategy for synthesizing argument positions — aggregate evidence, resolveVotes: 0GitHub stars: 417
- Argument VisualizationSOP for generating argument structure visualization — query graph forVotes: 0GitHub stars: 417
- Axis ExtractionTactic for systematically extracting axes of variation from literatureVotes: 0GitHub stars: 417
- Axis ValidationSOP for validating that candidate axes are independent and meaningful.Votes: 0GitHub stars: 417
- Causal Chain QuerySOP for tracing causal chains — follow edges from cause to effect throughVotes: 0GitHub stars: 417