All authors

Claude Skills by qhjqhj00
github.com/qhjqhj007,574 skills8 installs6,839 views
- Agentcaster Tornado Forecasting EvalEvaluates multimodal LLMs' ability to perform spatiotemporal reasoning and probabilistic risk forecasting for tornadoes by interactively querying weather data and generating geographic risk polygons. It measures forecasting accuracy, hallucination severity, and geometric precision against official meteorological baselines. Use when the user wants to benchmark on TornadoBench, or asks about evaluating this task. Reports TornadoBench.Votes: 0GitHub stars: 3
- Agentdrive Mcq EvalEvaluates large language models' context-sensitive reasoning and decision-making capabilities in autonomous driving scenarios. It probes physics-based calculations, policy compliance, risk interpretation, and maneuver optimization through multiple-choice questions derived from structured driving simulations. Use when the user wants to benchmark on AgentDrive-MCQ, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Agentds EvalThis benchmark evaluates AI agents and human-AI collaboration on domain-specific data science tasks across six industries. It probes the ability to perform feature engineering, integrate multimodal data (images, text, PDFs, JSON), and build predictive models that require genuine domain reasoning rather than generic pipelines. Use when the user wants to benchmark on AgentDS, or asks about evaluating this task. Reports quantile_score.Votes: 0GitHub stars: 3
- Agentehr EvalEvaluates autonomous clinical decision-making agents on Electronic Health Record (EHR) data. It probes multi-step reasoning, long-context dependency preservation, and robustness to distribution shifts across different hospital databases and clinical event types. Use when the user wants to benchmark on MIMIC-IV / MIMIC-III, or asks about evaluating this task. Reports average score.Votes: 0GitHub stars: 3
- Agentfuel EvalEvaluates LLM-based data analysis agents on their ability to execute domain-specific time-series queries, particularly focusing on stateful logic, temporal dependencies, and incident pattern detection. It probes whether agents can correctly interpret schemas, track state across sequential events, and identify anomalous behavior without relying on predefined time windows. Use when the user wants to benchmark on AgentFuel Benchmark, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Agentharm EvalThis benchmark evaluates the harmfulness and safety alignment of LLM-based agents by measuring their compliance with malicious, multi-step tasks that require coherent tool chaining. It probes whether models can be coerced into executing harmful behaviors through direct prompting or simple jailbreak templates, while tracking refusal rates and capability preservation. Use when the user wants to benchmark on AgentHarm, or asks about evaluating this task. Reports harm score.Votes: 0GitHub stars: 3
- Agenticcache EvalEvaluates the ability of embodied multi-agent systems to execute long-horizon, coordinated tasks efficiently using cache-driven asynchronous planning. It probes how well agents can reuse cached plan transitions to reduce LLM inference latency and token costs while maintaining high task success rates across diverse 3D simulation environments. Use when the user wants to benchmark on TDW-MAT, TDW-COOK, TDW-GAME, BEHAVIOR-1K, or asks about evaluating this task. Reports Success Rate.Votes: 0GitHub stars: 3
- AgentprmevalEvaluates LLM agents' ability to navigate simulated environments and execute multi-step plans to complete natural language instructions. It probes step-wise decision-making, goal proximity tracking, and sequential task execution across web shopping, grid-world navigation, and text-based crafting scenarios. Use when the user wants to benchmark on WebShop, BabyAI, TextCraft, or asks about evaluating this task. Reports success rate.Votes: 0GitHub stars: 3
- Agentquest EvalThis evaluation protocol measures LLM agent performance on multi-step reasoning tasks by tracking step-wise progress toward goal completion and the frequency of repetitive actions or states. It enables fine-grained debugging and architectural refinement beyond simple pass/fail success rates. Use when the user wants to benchmark on ALFWorld, Sudoku, or asks about evaluating this task. Reports progress rate.Votes: 0GitHub stars: 3
- Agentrecbench EvalThis benchmark evaluates LLM-based agentic recommender systems across three scenarios: classic, evolving-interest, and cold-start recommendation. It probes the agents' ability to dynamically plan, utilize textual interaction environments, and adapt to user preference shifts or data sparsity using structured user/item profiles and reviews. Use when the user wants to benchmark on Amazon, GoodReads, Yelp, or asks about evaluating this task. Reports Hit Rate@$N.Votes: 0GitHub stars: 3
- Agentrewardbench EvalThis benchmark evaluates the effectiveness of LLM-based judges in automatically assessing web agent trajectories. It probes the judges' ability to correctly predict task success, detect side effects, and identify repetitive actions by comparing their outputs against expert human annotations. Use when the user wants to benchmark on AgentRewardBench, or asks about evaluating this task. Reports precision.Votes: 0GitHub stars: 3
- Agentsafe EvalEvaluates the safety of embodied vision-language model agents when executing hazardous instructions in simulated indoor environments. It probes four key capabilities across perception, planning, and execution stages: object recognition accuracy, refusal to plan harmful actions, success in generating harmful plans, and success in physically executing them under both direct and jailbroken conditions. Use when the user wants to benchmark on AGENTSAFE, or asks about evaluating this task. Reports ...Votes: 0GitHub stars: 3
- Agentseval EvalProbes the clinical faithfulness, factual accuracy, and diagnostic logic of medical imaging report generation systems. It evaluates robustness to paraphrasing and semantic perturbations by decomposing assessment into interpretable reasoning stages that mimic radiologist workflows. Use when the user wants to benchmark on Five medical imaging datasets (names not provided in excerpt), or asks about evaluating this task. Reports AgentsEval Score.Votes: 0GitHub stars: 3
- Agentsynth EvalEvaluates the ability of multimodal language models to execute long-horizon, multi-step computer-use tasks on a desktop environment. It probes visual grounding, precise GUI interaction, state tracking, and error recovery across varying task complexities and software domains. Use when the user wants to benchmark on AgentSynth, or asks about evaluating this task. Reports success rate.Votes: 0GitHub stars: 3
- Agentvista EvalEvaluates the ability of multimodal agents to perform long-horizon, multi-step tool use in complex, realistic visual environments. It probes cross-image reasoning, constraint tracking, and robust grounding when interacting with dynamic tools like web search and code execution. Use when the user wants to benchmark on AgentVista, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Aghi Qa EvalEvaluates the perceptual quality and text-image correspondence of AI-generated human images, while also benchmarking the ability of models to identify visible and semantically distorted human body parts. Use when the user wants to benchmark on AGHI-QA, or asks about evaluating this task. Reports SRCC.Votes: 0GitHub stars: 3
- Agieval EvalThis benchmark evaluates foundation models on human-level cognitive abilities and general reasoning by testing them on a diverse collection of standardized admission and qualification exams. It probes domain-specific knowledge, analytical reasoning, and problem-solving across subjects like mathematics, law, logic, and languages. Use when the user wants to benchmark on AGIEval, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Agkphysics CccCompute agkphysics/ccc via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of agkphysics/ccc.Votes: 0GitHub stars: 3
- Agnn Citation EvalEvaluates semi-supervised node classification on citation networks using an attention-based graph neural network. Tests performance under fixed benchmark splits, random node sampling, and larger training sets to measure classification accuracy and attention interpretability. Use when the user wants to benchmark on CiteSeer, Cora, PubMed, or asks about evaluating this task. Reports classification accuracy.Votes: 0GitHub stars: 3
- Agri Met Recommendations EvalEvaluates the ability of LLMs to generate accurate and context-aware agricultural recommendations (sowing schedules, irrigation plans, risk mitigation) based on integrated weather, soil, and crop data. It specifically probes how multi-round prompt engineering improves recommendation quality compared to single-round and Chain-of-Thought baselines. Use when the user wants to benchmark on Agricultural Meteorological Dataset, or asks about evaluating this task. Reports Accuracy (Acc).Votes: 0GitHub stars: 3
- Agriculture Asr EvalEvaluates Automatic Speech Recognition (ASR) models on real-world agricultural field recordings across three Indian languages (Hindi, Telugu, Odia). It probes the models' ability to transcribe domain-specific terminology under challenging acoustic conditions like wind noise and multi-speaker overlap. Use when the user wants to benchmark on Agricultural Field Recordings, or asks about evaluating this task. Reports AWWER.Votes: 0GitHub stars: 3
- Agrigpt Omni EvalEvaluates a unified speech-vision-text model's capability in multilingual agricultural reasoning, covering text generation, vision-language QA, and multimodal speech understanding across open-ended and multiple-choice formats. Use when the user wants to benchmark on AgriBench-13K, AgriBench-VL-4K, AgriBench-Omni-2K, or asks about evaluating this task. Reports Accuracy.Votes: 0GitHub stars: 3
- Ahat Planning EvalEvaluates an agent's ability to decompose abstract, long-horizon household instructions into feasible, constraint-satisfying action plans. It probes intent inference, subgoal grounding, and robustness to environmental clutter and instruction ambiguity. Use when the user wants to benchmark on AHAT, Human Tasks, PARTNR, Behavior-1K, or asks about evaluating this task. Reports Success Rate (SR).Votes: 0GitHub stars: 3
- Ahnyeonchan Alignment And UniformityCompute ahnyeonchan/Alignment-and-Uniformity via the HuggingFace `evaluate` library. Use when the user has predictions + references and wants the canonical implementation of ahnyeonchan/Alignment-and-Uniformity.Votes: 0GitHub stars: 3
- Ahup 3d Pose EvalEvaluates monocular 3D human pose estimation models trained exclusively on synthetic 3D data and real 2D images, testing their ability to generalize to real-world 3D pose benchmarks without using any real 3D pose annotations during training. It probes domain adaptation capabilities, cross-dataset generalization, and the effectiveness of skeletal pose alignment strategies. Use when the user wants to benchmark on Human3.6M, MuPoTS, SURREAL, ScanAva+, MSCOCO, MPII Human Pose, or asks about evalu...Votes: 0GitHub stars: 3
- Ai Accelerator Training EvalEvaluates the computational performance and energy efficiency of various AI accelerators (CPUs, GPUs, TPUs) across standard deep learning workloads, including CNNs and NLP models. It measures how hardware architecture, numerical precision, and batch size impact training throughput and power consumption. Use when the user wants to benchmark on Standard DNN Workloads (ResNet50, Inception v3, Vgg16, LSTM, Deep Speech 2, Transformer), or asks about evaluating this task. Reports throughput.Votes: 0GitHub stars: 3
- Ai Benchmark EvalEvaluates the inference performance of mobile AI accelerators across major SoC vendors by running a standardized suite of deep learning models via TensorFlow Lite and NNAPI. It measures latency and accuracy to compare on-device AI capabilities against desktop hardware and track hardware evolution. Use when the user wants to benchmark on AI Benchmark 3.0, or asks about evaluating this task. Reports AI-Score.Votes: 0GitHub stars: 3
- Ai Face Fairness Bench EvalEvaluates the fairness and utility of AI-generated face detectors across demographic attributes (skin tone, gender, age) and intersectional groups. It measures how well detectors distinguish real vs. AI-generated faces while ensuring equitable performance across demographic subgroups. Use when the user wants to benchmark on AI-Face, or asks about evaluating this task. Reports $F_{MEO}$.Votes: 0GitHub stars: 3
- Ai Genbench EvalEvaluates the ability of AI-generated image detectors to generalize to novel, temporally subsequent generative models under realistic post-processing conditions. It measures how well detectors maintain performance when incrementally trained on historically ordered synthetic data and tested on unseen future generators. Use when the user wants to benchmark on AI-GenBench, or asks about evaluating this task. Reports AUROC, Accuracy.Votes: 0GitHub stars: 3
- Ai Paper Error Audit EvalEvaluates an LLM-based auditing system's ability to detect, categorize, and quantify objective mistakes in published AI research papers. It measures the system's precision against human verification and its recall against injected ground-truth errors across mathematical, textual, tabular, and cross-reference categories. Use when the user wants to benchmark on Published AI Papers (ICLR, NeurIPS, TMLR), or asks about evaluating this task. Reports precision.Votes: 0GitHub stars: 3
- Ai QualityEvaluates a feature-hierarchical edge inference framework's ability to dynamically allocate communication and computation resources to maximize AI quality under strict latency and energy constraints. Use when the user has predictions and gold and needs to compute AI quality (mAP).Votes: 0GitHub stars: 3
- Ai Red Teaming Ctf EvalEvaluates adversarial AI red-teaming capabilities by measuring participant success rates in bypassing LLM guardrails, manipulating model outputs, and extracting sensitive data through prompt injection and jailbreaking techniques. Use when the user wants to benchmark on AI Red Teaming CTF (CTF ID: 2604), or asks about evaluating this task. Reports solve_rate.Votes: 0GitHub stars: 3
- Ai Review Detection EvalEvaluates a style-based classifier's ability to detect AI-generated text in academic peer reviews and measures temporal generalization by tracking detection rates across consecutive years. Use when the user wants to benchmark on ICLR Peer Reviews, Nature Communications Peer Reviews, or asks about evaluating this task. Reports percentage_ai_detected.Votes: 0GitHub stars: 3
- Ai Writing Assistance EvalEvaluates how source disclosure and perceived AI authorship influence human editing behavior and subsequent peer-review acceptance decisions for scientific abstracts. Use when the user wants to benchmark on CS-Conference-Abstracts, or asks about evaluating this task. Reports accept/reject decision.Votes: 0GitHub stars: 3
- Ai4arctic Sod EvalThis benchmark evaluates pixel-wise sea ice stage of development (SOD) segmentation using dual-polarized SAR imagery. It probes a model's ability to accurately classify ice types under varying quantization levels and measures hardware efficiency across different computing platforms. Use when the user wants to benchmark on AI4Arctic Sea Ice Dataset, or asks about evaluating this task. Reports F1 score.Votes: 0GitHub stars: 3
- Ai4skin Subtyping EvalEvaluates histopathology foundation models' ability to extract center-invariant, biologically relevant features for skin cancer subtyping. It measures representation bias toward scanning centers and downstream classification performance under multiple instance learning frameworks. Use when the user wants to benchmark on AI4SkIN, or asks about evaluating this task. Reports Balanced Accuracy (BACC).Votes: 0GitHub stars: 3
- Aibench EvalProbes the end-to-end latency and micro-architectural efficiency of AI-accelerated internet service workloads. It measures how AI components impact service latency and GPU execution stalls during both online inference and offline training. Use when the user wants to benchmark on AIBench E-commerce Search Workload, or asks about evaluating this task. Reports Latency (avg, p90, p99).Votes: 0GitHub stars: 3
- Aibench Scenario EvalEvaluates the end-to-end system-level performance and tail latency of AI-driven online services by simulating real-world user workloads. It probes how cascading interactions between AI and non-AI components affect overall service quality, and tests the validity of statistical queueing models for predicting system latency. Use when the user wants to benchmark on AIBench Scenario (E-commerce & Translation Intelligence), or asks about evaluating this task. Reports latency (avg, p90, p99).Votes: 0GitHub stars: 3
- Aibench Training EvalEvaluates AI training workloads by measuring model complexity, computational cost, convergence rate, and micro-architectural behavior to assess benchmark diversity, repeatability, and cost against industry standards like MLPerf. Use when the user wants to benchmark on AIBench Training, or asks about evaluating this task. Reports convergent_rate.Votes: 0GitHub stars: 3
- Aicabench EvalEvaluates Vision-Language Models on affective image content analysis across three dimensions: Emotion Understanding (identifying emotions in images), Emotion Reasoning (inferring emotional causes/context), and Emotion-Guided Content Generation (producing text guided by emotional intent). It probes models' ability to perceive, reason about, and generate content based on visual emotional cues, including sensitivity to abstract art and reliance on facial shortcuts. Use when the user wants to ben...Votes: 0GitHub stars: 3
- Aid Aerial Scene Classification EvalEvaluates the ability of computer vision models to classify aerial imagery into distinct scene categories. It probes robustness to high intra-class diversity and low inter-class similarity in remote sensing data. Use when the user wants to benchmark on AID, or asks about evaluating this task. Reports accuracy.Votes: 0GitHub stars: 3
- Aidabench EvalEvaluates LLM-driven document analysis agents on end-to-end data analytics workflows, including question answering, data visualization, and file generation. It probes multi-step numerical reasoning, cross-data consistency, and long-horizon planning on heterogeneous real-world documents. Use when the user wants to benchmark on AIDABench, or asks about evaluating this task. Reports Pass@3.Votes: 0GitHub stars: 3
- Aide Benchmark EvalEvaluates the performance of LLMs fine-tuned on synthetic data generated by AIDE across a suite of standard knowledge and reasoning benchmarks. It probes zero-shot and few-shot generalization capabilities compared to models fine-tuned on human-curated gold data. Use when the user wants to benchmark on MMLU, FinBen, ARC-Challenge, GSM8K, TruthfulQA, MedQA, BIG-Bench, or asks about evaluating this task. Reports zero-shot accuracy.Votes: 0GitHub stars: 3
- Aidovecl EvalEvaluates the effectiveness of AI-generated outpainted vehicle images as data augmentation for training object detection models. It probes the model's ability to generalize to real-world vehicle classification and bounding box localization when trained on synthetically augmented data. Use when the user wants to benchmark on AIDOVECL augmented dataset, or asks about evaluating this task. Reports F1 Score.Votes: 0GitHub stars: 3
- Aifl Streamflow Forecast EvalEvaluates a model's ability to forecast daily specific streamflow at global gauging stations under temporal generalization. It specifically probes robustness to domain shifts between reanalysis pre-training data and operational forecast fine-tuning data, testing whether the model maintains performance when transitioning from historical reanalysis to real-time operational forcing. Use when the user wants to benchmark on CARAVAN v1.5, or asks about evaluating this task. Reports KGE.Votes: 0GitHub stars: 3
- Aigc Detection Accuracy EvalEvaluates the cross-generator generalization capability of AI-generated image (AIGC) detectors. It probes whether models trained on a specific generator (SDv1.4) can accurately distinguish real from fake images produced by diverse, unseen generative models and in-the-wild sources. Use when the user wants to benchmark on GenImage, GenImage++, Chameleon, or asks about evaluating this task. Reports Accuracy (ACC).Votes: 0GitHub stars: 3
- Aigcbench EvalEvaluates the performance of image-to-video (I2V) generation models across multiple quality and alignment dimensions. It probes how well models preserve input image fidelity, generate coherent motion, align with text prompts, maintain temporal consistency, and produce high-quality video output. Use when the user wants to benchmark on AIGCBench Dataset, or asks about evaluating this task. Reports video quality.Votes: 0GitHub stars: 3
- Aigi Detection EvalEvaluates the ability of models to distinguish real photographs from AI-generated images across diverse, out-of-distribution, and post-processed scenarios. It probes both low-level pixel artifact detection and high-level semantic consistency checking to measure real-world generalization. Use when the user wants to benchmark on Chameleon, WildRF, AIGI-Bench, Co-SPY-Bench (in-the-wild), BFree-Online, AIGI-Now, GenImage, DRCT-2M, AIGCDetectBenchmark, or asks about evaluating this task. Reports B...Votes: 0GitHub stars: 3
- Aigibench EvalEvaluates the generalization, robustness to image degradation, and sensitivity to data augmentation and pre-processing of AI-generated image (AIGI) detectors across 25 diverse test datasets spanning GANs, diffusion models, and face-swap/manipulation methods. Use when the user wants to benchmark on AIGIBench, or asks about evaluating this task. Reports F.Acc..Votes: 0GitHub stars: 3
- Aigiq 20k EvalThis benchmark evaluates the perceptual quality and text-to-image alignment of AI-generated images. It benchmarks objective quality assessment models against large-scale human subjective ratings to measure how well automated metrics correlate with human perception. Use when the user wants to benchmark on AIGIQA-20K, or asks about evaluating this task. Reports SRoCC.Votes: 0GitHub stars: 3