All authors

Claude Skills by qhjqhj00
github.com/qhjqhj007,574 skills8 installs6,839 views
- Aigvdbench EvalEvaluates the ability of AI-generated video detectors to distinguish between real and synthetically generated videos across diverse generation models, tasks (T2V, I2V, V2V), and temporal/spatial artifacts. Use when the user wants to benchmark on AIGVDBench, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Aiml Tuda IsomorphicperturbationtestingCompute AIML-TUDA/IsomorphicPerturbationTesting via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of AIML-TUDA/IsomorphicPerturbationTesting.Votes: 0GitHub stars: 3
- Aiml Tuda VerifiablerewardsforscalablelogicalreasoningCompute AIML-TUDA/VerifiableRewardsForScalableLogicalReasoning via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of AIML-TUDA/VerifiableRewardsForScalableLogicalReasoning.Votes: 0GitHub stars: 3
- Aiotbench EvalEvaluates AI inference performance across diverse image classification model architectures on mobile and embedded devices. It measures the trade-off between inference speed and computational efficiency to compare models, frameworks, and hardware. Use when the user wants to benchmark on ImageNet 2012, or asks about evaluating this task. Reports VIPS.Votes: 0GitHub stars: 3
- Aipperf EvalEvaluates the end-to-end performance and weak scalability of heterogeneous AI-HPC systems using AutoML workloads. It measures how efficiently clusters execute dynamically scaling machine learning training and inference tasks across varying numbers of nodes. Use when the user wants to benchmark on CIFAR10, or asks about evaluating this task. Reports cumulative OPS.Votes: 0GitHub stars: 3
- Air Bench EvalEvaluates Large Audio-Language Models on foundational audio comprehension across speech, natural sounds, and music, as well as open-ended instruction-following via generative responses. It probes the model's ability to understand mixed audio, follow complex prompts, and produce accurate, contextually relevant text. Use when the user wants to benchmark on AIR-Bench, or asks about evaluating this task. Reports GPT-4 alignment strategy.Votes: 0GitHub stars: 3
- Air Quality Forecasting EvalEvaluates a regression model's ability to forecast hyper-local air pollutant concentrations using fine-grained traffic intensity descriptors. It probes how well traffic patterns across different spatial rings and colors correlate with specific pollutant levels under varying training station configurations. Use when the user wants to benchmark on Mexico City Traffic & Pollution Dataset, or asks about evaluating this task. Reports RMSE.Votes: 0GitHub stars: 3
- Airs Bench EvalEvaluates AI research agents across the full scientific lifecycle, including idea generation, experiment design, and iterative refinement. Agents must generate and execute code to train models on specified datasets without baseline code, testing reasoning, generalization, and solution exploration capabilities. Use when the user wants to benchmark on AIRS-Bench, or asks about evaluating this task. Reports average normalized score.Votes: 0GitHub stars: 3
- Airscape 6dof EvalEvaluates a generative world model's ability to predict first-person future video observations under specified 6DoF aerial motion intentions. It probes spatio-temporal consistency, motion alignment, and counterfactual reasoning in 3D aerial environments. Use when the user wants to benchmark on AirScape Dataset, or asks about evaluating this task. Reports IAR.Votes: 0GitHub stars: 3
- Aishell3 Tts EvalEvaluates the ability of a multi-speaker TTS system to synthesize high-fidelity Mandarin speech that preserves speaker identity across both seen and unseen speakers. It probes zero-shot voice cloning capability and generalization to novel speakers using objective speaker verification metrics. Use when the user wants to benchmark on AISHELL-3, or asks about evaluating this task. Reports SV-EER.Votes: 0GitHub stars: 3
- Aist Dance EvalEvaluates the quality, diversity, and music-motion synchronization of generated 3D dance sequences. It probes a model's ability to synthesize physically plausible, choreographically diverse, and rhythm-aligned human motion from audio input. Use when the user wants to benchmark on AIST++, or asks about evaluating this task. Reports FID_k.Votes: 0GitHub stars: 3
- Aitw EvalEvaluates an agent's ability to infer and execute multi-step visual actions on Android devices from natural language instructions. It specifically probes Out-of-Distribution generalization across unseen Android OS versions, instruction language patterns (subjects/verbs), and app/web domains. Use when the user wants to benchmark on AITW, or asks about evaluating this task. Reports average score.Votes: 0GitHub stars: 3
- Aj Jaggedness PenaltyProposes a diagnostic metric to quantify the discrepancy between macro-averaged benchmark scores and experienced reliability in deployment. It accounts for uneven task coverage by measuring the dispersion of error rates across domains, highlighting how gap-uniform evaluation can misstate actual user experience. Use when the user has predictions and gold and needs to compute CV_d(e_d).Votes: 0GitHub stars: 3
- Akki2825 Accents Unplugged EvalCompute akki2825/accents_unplugged_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of akki2825/accents_unplugged_eval.Votes: 0GitHub stars: 3
- Alagin Vc EvalEvaluates voice conversion systems on speech quality and speaker similarity using subjective human ratings. It probes the ability of models to convert speech between speakers (specifically inter-gender) while preserving linguistic content and target speaker identity. Use when the user wants to benchmark on ALAGIN Japanese Speech Database Set B, or asks about evaluating this task. Reports Mean Opinion Score (MOS) for Speech Quality.Votes: 0GitHub stars: 3
- Albertgong1 My MetricCompute albertgong1/my_metric via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of albertgong1/my_metric.Votes: 0GitHub stars: 3
- Alden Vrdu EvalEvaluates vision-language models' ability to actively navigate long, visually rich documents to gather evidence and answer complex queries. It probes multi-turn reasoning, retrieval accuracy, and the effectiveness of direct page-index access versus semantic search. Use when the user wants to benchmark on MMLongBench, LongDocURL, PaperTab, PaperText, FetaTab, DUDE-sub, or asks about evaluating this task. Reports GPT-4o–judged answer accuracy (Acc).Votes: 0GitHub stars: 3
- Ale 60 Games EvalThis benchmark evaluates reinforcement learning agents across 60 Atari 2600 games using a stochastic environment variant with sticky actions. It measures sample efficiency and learning stability by tracking performance at multiple frame-count thresholds (10M to 200M). Use when the user wants to benchmark on Arcade Learning Environment (ALE), or asks about evaluating this task. Reports score averages.Votes: 0GitHub stars: 3
- Alephbert EvalEvaluates pre-trained Hebrew language models on core NLP tasks including morphological analysis, named entity recognition, and sentiment analysis. It measures how well the models handle Hebrew-specific linguistic features and resource-scarce language challenges compared to existing baselines. Use when the user wants to benchmark on SPMRL Hebrew Section, UD treebanks Hebrew Section, Ben-Mordecai and Elhadad corpus, NEMO corpus, Amram et al. (2018) corpus (cleaned), or asks about evaluating thi...Votes: 0GitHub stars: 3
- Alfred EvalEmbodied instruction following in a simulated household environment, requiring an agent to execute long-horizon navigation and object manipulation tasks based on natural language commands. Use when the user wants to benchmark on ALFRED, or asks about evaluating this task. Reports Success Rate (SR).Votes: 0GitHub stars: 3
- Algerian Dialect EvalEvaluates cross-lingual and cross-script transfer performance for sentiment analysis and topic classification on a novel multi-layer Algerian dialect corpus. Probes how script differences (Latin/NArabizi vs. Arabic/Persian/Urdu) and typological similarity impact classification accuracy in code-switched, under-resourced vernaculars. Use when the user wants to benchmark on Algerian Dialect Corpus (NArabizi), or asks about evaluating this task. Reports Macro F1.Votes: 0GitHub stars: 3
- Algonauts 2019 EvalThis benchmark evaluates a model's ability to predict human visual brain activity during object recognition. It compares model representations against fMRI and MEG neural recordings using representational similarity analysis (RSA) across spatial (EVC vs IT) and temporal (early vs late processing) dimensions. Use when the user wants to benchmark on Algonauts 2019 Challenge, or asks about evaluating this task. Reports noise-normalized variance explained.Votes: 0GitHub stars: 3
- Alhitawimohammed22 Cer Hu Evaluation MetricsCompute AlhitawiMohammed22/CER_Hu-Evaluation-Metrics via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of AlhitawiMohammed22/CER_Hu-Evaluation-Metrics.Votes: 0GitHub stars: 3
- Align Vl EvalEvaluates a vision-language model's cross-modal retrieval capabilities (matching images to text and vice versa) and zero-shot image classification performance without task-specific fine-tuning. It also measures transfer learning effectiveness on downstream visual benchmarks via linear probing and full fine-tuning. Use when the user wants to benchmark on Flickr30K, MSCOCO, or asks about evaluating this task. Reports R@10.Votes: 0GitHub stars: 3
- Alignment Research Classifier EvalEvaluates a logistic regression model on SPECTER embeddings to distinguish AI alignment research articles from adjacent research on arXiv. Probes the model's ability to capture domain-specific semantic patterns and citation-driven textual features for automated literature filtering. Use when the user wants to benchmark on arXiv Alignment Research Corpus, or asks about evaluating this task. Reports AUC.Votes: 0GitHub stars: 3
- Alloprof Ir EvalEvaluates information retrieval capabilities in an educational context by testing a model's ability to retrieve relevant reference pages or similar past questions given a student's query. It probes handling of noisy text (spelling/grammar errors), multimodal inputs (images, formulas), and grade-aware language complexity. Use when the user wants to benchmark on Alloprof, or asks about evaluating this task. Reports nDCG.Votes: 0GitHub stars: 3
- Alloy Phase Diagram ValidationEvaluates a Wang-Landau sampling method combined with cluster expansion for predicting thermodynamic phase diagrams of binary alloys. It probes the method's ability to capture ordering and phase-separation tendencies, and accurately reproduce experimental phase boundaries and transition temperatures. Use when the user wants to benchmark on Cu-Au alloy, Pd-Rh alloy, or asks about evaluating this task. Reports cross-validation score.Votes: 0GitHub stars: 3
- Allvb EvalEvaluates multimodal large language models' ability to comprehend hour-long videos across nine distinct tasks, including classification, recognition, localization, captioning, emotion recognition, and needle-in-a-haystack retrieval. It specifically probes temporal reasoning, detail extraction, and long-context retention over extended video durations. Use when the user wants to benchmark on ALLVB, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Alm Bench EvalThis benchmark evaluates the cultural and linguistic reasoning capabilities of large multimodal models across 100 languages. It probes visual understanding and cultural knowledge through generic and culturally specific domains, testing both closed-form and open-ended question answering. Use when the user wants to benchmark on ALM-bench, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Almanacs EvalEvaluates whether language model explanations (e.g., weights, qualitative descriptions) enable a second predictor model to accurately simulate and predict the behavior of a synthetic linear model across safety-relevant scenarios. The benchmark specifically probes simulatability and robustness to distributional shift between training and test variable values. Use when the user wants to benchmark on ALMANACS Synthetic Dataset, or asks about evaluating this task. Reports probability.Votes: 0GitHub stars: 3
- Aloha Bimanual Manipulation EvalThis evaluation probes a robot policy's ability to perform precise, temporally extended bimanual manipulation tasks using only visual and proprioceptive inputs. It specifically tests the model's robustness to compounding errors, non-Markovian dynamics, and perception challenges like transparent or low-contrast objects. Use when the user wants to benchmark on ALOHA Fine Manipulation Tasks, or asks about evaluating this task. Reports success rate.Votes: 0GitHub stars: 3
- Alope Qe EvalThis evaluation probes the ability of LLM-based frameworks to estimate translation quality in a reference-free setting by predicting continuous quality scores for source-target sentence pairs across multiple low-resource language directions. It specifically tests how intermediate Transformer layer representations and adaptive regression heads improve cross-lingual alignment and quality prediction compared to standard fine-tuning or zero-shot prompting. Use when the user wants to benchmark on ...Votes: 0GitHub stars: 3
- Alora Peft EvalEvaluates the effectiveness of dynamic low-rank adaptation (LoRA) for fine-tuning large language models across classification, question answering, and instruction generation tasks. It measures how well rank allocation strategies preserve performance while maintaining or reducing tunable parameter counts. Use when the user wants to benchmark on SQuAD, BoolQ, COPA, ReCoRD, SST-2, RTE, QNLI, Alpaca, MT-Bench, E2E, or asks about evaluating this task. Reports accuracy, GPT-4 score.Votes: 0GitHub stars: 3
- Alpaca Eval Lc Winrate EvalEvaluates the alignment quality of language models by measuring their win rate against a baseline on the AlpacaEval benchmark. It specifically uses length-controlled (LC) win rates to mitigate the known bias toward longer model outputs in standard auto-annotator evaluations. Use when the user wants to benchmark on alpaca_eval, or asks about evaluating this task. Reports AlpacaEval length-controlled (LC) win rate.Votes: 0GitHub stars: 3
- Alpacaeval EvalEvaluates LLM response quality via pairwise win rates against a baseline, while specifically probing the metric's susceptibility to length bias, gameability via verbosity prompting, and robustness to adversarial truncation. Use when the user wants to benchmark on AlpacaEval, or asks about evaluating this task. Reports Win rate.Votes: 0GitHub stars: 3
- Alpbench EvalEvaluates active learning pipelines by comparing query strategies paired with tabular classifiers across multiple datasets. It measures how efficiently pipelines improve test performance as the labeled data budget increases, highlighting the interplay between learner choice and query strategy. Use when the user wants to benchmark on OpenML-CC18 and TabZilla Benchmark Suite, or asks about evaluating this task. Reports AUBC (Area Under the Budget Curve).Votes: 0GitHub stars: 3
- Alphacode EvalEvaluates a model's ability to generate correct, executable code for competitive programming problems under strict submission limits. It probes algorithmic reasoning, code synthesis, and the capacity to pass hidden test cases after filtering on provided examples. Use when the user wants to benchmark on CodeContests, or asks about evaluating this task. Reports solve rate.Votes: 0GitHub stars: 3
- Alpharesearch Algo Discovery EvalEvaluates an LLM-based autonomous agent's ability to discover novel algorithms through iterative idea generation, code modification, and execution-based verification. It probes the model's capacity for scientific reasoning, program synthesis, and optimization under simulated peer-review feedback. Use when the user wants to benchmark on AlphaResearch Algorithm Discovery Problems, or asks about evaluating this task. Reports win_rate (excel@best > 0).Votes: 0GitHub stars: 3
- Alpsbench EvalAlpsBench evaluates the full lifecycle of LLM personalization, including extracting structured memories from dialogue, dynamically updating them, retrieving relevant memories under distractors, and utilizing them to generate aligned responses across dimensions like persona awareness, preference following, and emotional intelligence. Use when the user wants to benchmark on AlpsBench, or asks about evaluating this task. Reports F1 score (exact match).Votes: 0GitHub stars: 3
- Alrm Manipulation EvalEvaluates an agentic LLM's ability to plan and execute multistep robotic manipulation tasks in simulation. It tests closed-loop reasoning via ReAct-style loops, comparing code-as-policy and tool-as-policy execution modes across linguistically diverse tasks. Use when the user wants to benchmark on ALRM Simulation Benchmark, or asks about evaluating this task. Reports task_completion.Votes: 0GitHub stars: 3
- Alvinasvk Accents Unplugged EvalCompute alvinasvk/accents_unplugged_eval via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of alvinasvk/accents_unplugged_eval.Votes: 0GitHub stars: 3
- Alzheimer Mri 4class EvalEvaluates multi-class classification performance on Alzheimer's disease MRI scans to assess a model's ability to distinguish between different stages of dementia and healthy controls under resource-constrained hardware conditions. Use when the user wants to benchmark on Alzheimer MRI 4 Classes Dataset, or asks about evaluating this task. Reports Accuracy.Votes: 0GitHub stars: 3
- Am Lora Cl EvalEvaluates a model's ability to continuously learn multiple text classification tasks without catastrophic forgetting, measuring how well it retains knowledge of previous tasks while adapting to new ones. Use when the user wants to benchmark on Standard CL benchmarks, Large number of tasks benchmark, or asks about evaluating this task. Reports average results.Votes: 0GitHub stars: 3
- Amazon M2 EvalEvaluates session-based recommendation and text generation capabilities across multiple languages and locales. It probes a model's ability to predict the next product in a shopping session, transfer knowledge across domain-shifted locales, and generate product titles from session context. Use when the user wants to benchmark on Amazon-M2, or asks about evaluating this task. Reports next-product prediction.Votes: 0GitHub stars: 3
- Amazon Review Summarization EvalEvaluates the ability of abstractive summarization models to generate concise, informative summaries of product reviews while preserving aspect and opinion details. It measures lexical overlap with human-written reference summaries. Use when the user wants to benchmark on Amazon Reviews (Healthcare & Electronics), or asks about evaluating this task. Reports ROUGE-1.Votes: 0GitHub stars: 3
- Amazon Stark Skb EvalEvaluates the ability of neural retriever-reranker pipelines to accurately retrieve relevant product entities from semi-structured e-commerce knowledge graphs using natural language queries. It probes semantic matching, cross-encoder reranking effectiveness, and the impact of graph-based augmentation on retrieval precision and recall. Use when the user wants to benchmark on Amazon STaRK SKB, or asks about evaluating this task. Reports Hit@1.Votes: 0GitHub stars: 3
- Ambigqa EvalThis benchmark evaluates a model's ability to identify ambiguous open-domain questions, generate multiple plausible answer spans, and produce disambiguated question rewrites that distinguish between different interpretations of the same query. Use when the user wants to benchmark on AMBIGNQ, or asks about evaluating this task. Reports F1ans.Votes: 0GitHub stars: 3
- Ambiguous Emotion Recognition EvalEvaluates audio-language models' ability to recognize ambiguous emotions in speech by predicting full emotion probability distributions and dominant class labels. It specifically probes how test-time scaling (TTS) strategies and model capacity interact with varying levels of emotional ambiguity to improve or degrade recognition performance. Use when the user wants to benchmark on IEMOCAP, MSP-Podcast, CREMA-D, or asks about evaluating this task. Reports JS divergence.Votes: 0GitHub stars: 3
- Ambiqt EvalEvaluates a model's ability to generate diverse, semantically valid SQL queries when a natural language question is ambiguous. It specifically probes whether the model can cover multiple valid interpretations of the same query within a limited set of top-k outputs. Use when the user wants to benchmark on AmbiQT, SPIDER, Kaggle DBQA, or asks about evaluating this task. Reports BothInTopK.Votes: 0GitHub stars: 3
- Ambisql EvalEvaluates a Text-to-SQL system's ability to generate correct SQL from ambiguous natural language queries when integrated with an interactive ambiguity resolution module. It also measures the system's precision, recall, and F1 in detecting and classifying specific types of schema-mapping and reasoning ambiguities. Use when the user wants to benchmark on AmbiSQL Constructed Dataset, or asks about evaluating this task. Reports Exact Match accuracy.Votes: 0GitHub stars: 3