All authors

Claude Skills by feiyang-k
github.com/feiyang-k258 skills0 installs459 views
- Protocols> Variant of `BENCHMARK.md`. The agent is given exactly one of `protocols/plain.md`, `protocols/instruction.md`, or `protocols/skill.md` at launch, depending on the chosen launch profile. The skill-grounded variant additionally exposes a library of paper-derived skill cards mounted at `/workspace/skills/<category>/<paper-slug>/SKILL.md` inside the container — agents are required to read a quota of them before editing curation code. The host-side source of those cards is the repository's top-l...Votes: 0GitHub stars: 9
- A Okvqa A Benchmark For Visual Question Answering Using World Knowledge Arxiv 2206 01718v2Use this skill when you need a VQA dataset requiring world knowledge and commonsense reasoning beyond what is visible in the image. Avoid it when your VQA task only requires visual perception without external knowledge.Votes: 0GitHub stars: 9
- Activitynet A Large Scale Video Benchmark For Human Activity Understanding Arxiv 1503 02001v3Use this skill when you need a large-scale video activity understanding dataset with temporal annotations for training video-language models. Avoid it when short-clip action recognition is sufficient.Votes: 0GitHub stars: 9
- Asr Whisper For Video Transcription Arxiv 2212 04356v1Use this skill when you want to use Whisper ASR to transcribe speech from videos for creating video-text training data. Avoid it when your videos do not have speech or you have existing transcripts.Votes: 0GitHub stars: 9
- Biomedclip A Multimodal Biomedical Foundation Model Pretrained From Fifteen Million Scientific Image Text Pairs Arxiv 2303 00915v4Use this skill when you want to build a biomedical CLIP model trained on 15M scientific image-text pairs from PMC. Avoid it when you need a general-purpose CLIP model or lack access to PMC data.Votes: 0GitHub stars: 9
- Blink Multimodal Large Language Models Can See But Not Perceive Arxiv 2404 12390v2Use this skill when you need a benchmark revealing perception gaps in VLMs — tasks that are trivially easy for humans but hard for models. Avoid it when standard VLM benchmarks show sufficient performance.Votes: 0GitHub stars: 9
- Cambrian A Data Centric Benchmark For Multimodal Ai Arxiv Cambrian Bench 2024Use this skill when you need a vision-centric evaluation benchmark that tests visual perception rather than language ability in VLMs. Avoid it when standard VLM benchmarks are sufficient.Votes: 0GitHub stars: 9
- Cc12m Conceptual 12m Pushing Web Scale Image Text Pre Training To Recognize Long Tail Visual Concepts Arxiv 2102 08981v2Use this skill when you want to scale up an image-text dataset from 3M to 12M by relaxing filtering criteria to capture long-tail visual concepts. Avoid it when you need very clean data and cannot tolerate increased noise from relaxed filtering.Votes: 0GitHub stars: 9
- Chartqa A Benchmark For Question Answering About Charts With Visual And Logical Reasoning Arxiv 2203 10244v1Use this skill when you need a QA dataset about charts requiring both visual perception and logical reasoning. Avoid it when you need VQA for natural images rather than charts and graphs.Votes: 0GitHub stars: 9
- Clip Benchmark A Benchmark Suite For Evaluation Of Vision Language Models Github Clip Benchmark 2023Use this skill when you need a standardized benchmark suite for evaluating CLIP-style vision-language encoders across dozens of datasets. Avoid it when you only evaluate on ImageNet zero-shot.Votes: 0GitHub stars: 9
- Coco Captions Constructing A Large Scale Dataset For Image Description Arxiv Coco Captions 2015Use this skill when you need gold-standard image captions collected through crowdsourcing for VLM training and evaluation. Avoid it when web-crawled captions are sufficient or you have your own caption data.Votes: 0GitHub stars: 9
- Coco Microsoft Coco Common Objects In Context Arxiv 1405 0312v3Use this skill when you need a foundational multi-purpose vision dataset with captions, detections, segmentation, and keypoints for VLM development. Avoid it when you need web-scale data rather than a carefully annotated 330K image dataset.Votes: 0GitHub stars: 9
- Conceptual Captions A Cleaned Hypernymed Image Alt Text Dataset For Automatic Image Captioning Arxiv 1809 00470v1Use this skill when you want to build an image captioning dataset from web alt-text using automated cleaning and hypernym replacement. Avoid it when you need fine-grained or domain-specific captions that web alt-text cannot provide.Votes: 0GitHub stars: 9
- Coyo 700m Image Text Pair Dataset Github Kakaobrain Coyo 700mUse this skill when you need a large-scale open image-text dataset from Common Crawl as an alternative to LAION with transparent collection methodology. Avoid it when LAION-5B or other existing datasets are sufficient for your needs.Votes: 0GitHub stars: 9
- Docvqa A Dataset For Vqa On Document Images Arxiv 2007 00398v3Use this skill when you need a VQA dataset on document images for training and evaluating document understanding VLMs. Avoid it when you need scene-level VQA rather than document understanding.Votes: 0GitHub stars: 9
- Ego4d Around The World In 3000 Hours Of Egocentric Video Arxiv 2110 07058v2Use this skill when you need a massive egocentric video dataset with diverse annotations for training video-language models. Avoid it when you do not work with egocentric/first-person video or need third-person video data.Votes: 0GitHub stars: 9
- Emu Generative Pretraining In Multimodality Arxiv 2307 05222v2Use this skill when you want to curate diverse interleaved multimodal data combining web pages, image-text pairs, and video for generative pretraining. Avoid it when you only need paired image-text data without interleaved document structure.Votes: 0GitHub stars: 9
- Flamingo A Visual Language Model For Few Shot Learning Arxiv 2204 14198v2Use this skill when you need to collect and train on interleaved image-text web data for few-shot visual learning. Avoid it when you only have paired image-caption data without interleaved document structure.Votes: 0GitHub stars: 9
- Flickr30k Entities Collecting Region To Phrase Correspondences For Richer Image To Sentence Models Arxiv 1505 04870v3Use this skill when you need a dataset linking noun phrases in captions to bounding box regions in images for visual grounding. Avoid it when you do not need phrase-level grounding or the 31K image scale is too small.Votes: 0GitHub stars: 9
- Glamm An Open Source Ai Framework For Data Processing Arxiv Img2dataset 2023Use this skill when you need to efficiently download billions of images from URLs for building large-scale image-text datasets. Avoid it when you already have images locally or only need a small number.Votes: 0GitHub stars: 9
- Grit General Robust Image Task Benchmark Arxiv 2306 14818v2Use this skill when you need a benchmark testing VLM robustness across multiple visual tasks with distribution shifts. Avoid it when standard in-distribution benchmarks are sufficient.Votes: 0GitHub stars: 9
- Hallusionbench An Advanced Diagnostic Suite For Entangled Language Hallucination And Visual Illusion In Large Vision Lan Arxiv 2310 14566v3Use this skill when you need a diagnostic benchmark testing VLM vulnerability to both language hallucination and visual illusion. Avoid it when you only need standard hallucination evaluation (use POPE instead).Votes: 0GitHub stars: 9
- Howto100m Learning A Text Video Embedding By Watching Hundred Million Narrated Video Clips Arxiv 1906 02604v2Use this skill when you want to mine narrated instructional videos from YouTube as video-text training data at hundred-million scale. Avoid it when you need clean video captions rather than noisy ASR transcripts.Votes: 0GitHub stars: 9
- Imagenet A Large Scale Hierarchical Image Database Crossref Imagenet 2009Use this skill when you need a large-scale hierarchical image dataset organized by WordNet synsets for vision model pretraining and evaluation. Avoid it when you need image-text pairs rather than image-label data.Votes: 0GitHub stars: 9
- Internvid A Large Scale Video Text Dataset For Multimodal Understanding And Generation Arxiv 2307 06942v3Use this skill when you need a 234M video-text dataset constructed with multi-scale captioning from ASR, metadata, and generated descriptions. Avoid it when you only need image-text data or a small video dataset.Votes: 0GitHub stars: 9
- Kinetics 400 A Large Video Understanding Dataset Arxiv 1705 06950v1Use this skill when you need a large-scale action recognition dataset with 400 categories and 300K+ video clips for video encoder pretraining. Avoid it when temporal reasoning rather than action recognition is your focus.Votes: 0GitHub stars: 9
- Laion 5b An Open Large Scale Dataset For Training Next Generation Image Text Models Arxiv 2210 08402v1Use this skill when you need to build a multi-billion scale open image-text dataset from Common Crawl using CLIP score filtering. Avoid it when you need a small curated dataset or cannot handle the infrastructure for billions of image downloads.Votes: 0GitHub stars: 9
- Languagebind Extending Video Language Pretraining To N Modality By Language Based Semantic Alignment Arxiv 2310 01852v4Use this skill when you want to align multiple modalities (video, audio, depth, thermal, infrared) to language for unified multi-modal understanding. Avoid it when you only need image-language alignment without other modalities.Votes: 0GitHub stars: 9
- Learning Transferable Visual Models From Natural Language Supervision Arxiv 2103 00020v1Use this skill when you need to construct a large-scale image-text dataset from the web for contrastive vision-language pretraining. Avoid it when you already have curated paired data or need fine-grained region-level annotations.Votes: 0GitHub stars: 9
- Lmms Eval Reality Check On The Evaluation Of Large Multimodal Models Arxiv 2407 12772v3Use this skill when you need a unified evaluation framework for consistently evaluating VLMs across dozens of benchmarks. Avoid it when you only evaluate on 1-2 benchmarks.Votes: 0GitHub stars: 9
- Lvis A Dataset For Large Vocabulary Instance Segmentation Arxiv 1908 03195v2Use this skill when you need a large-vocabulary instance segmentation dataset with 1,200+ categories handling the long tail of visual concepts. Avoid it when COCO's 80 categories are sufficient for your needs.Votes: 0GitHub stars: 9
- Mathvista Evaluating Mathematical Reasoning Of Foundation Models In Visual Contexts Arxiv 2310 02255v4Use this skill when you need a benchmark testing mathematical reasoning in visual contexts including charts, plots, diagrams, and geometry. Avoid it when you do not need mathematical reasoning evaluation.Votes: 0GitHub stars: 9
- Mimic Cxr A De Identified Publicly Available Database Of Chest Radiographs With Free Text Reports Arxiv Mimic Cxr 2019Use this skill when you need a large-scale medical image-report dataset for training biomedical VLMs on chest X-ray understanding. Avoid it when you do not work with medical imaging.Votes: 0GitHub stars: 9
- Mm Vet Evaluating Large Multimodal Models For Integrated Capabilities Arxiv 2308 02490v3Use this skill when you need a benchmark evaluating VLMs on tasks requiring integrated use of multiple capabilities simultaneously. Avoid it when individual capability testing is sufficient.Votes: 0GitHub stars: 9
- Mmbench Is Your Multi Modal Model An All Around Player Arxiv 2307 06281v4Use this skill when you need a multi-dimensional benchmark evaluating VLMs across 20+ ability dimensions with robust circular evaluation. Avoid it when you only need a single-dimension evaluation or have simpler benchmarks.Votes: 0GitHub stars: 9
- Mme A Comprehensive Evaluation Benchmark For Multimodal Large Language Models Arxiv 2306 13394v2Use this skill when you need a comprehensive VLM benchmark testing both perception and cognition abilities with yes/no questions. Avoid it when you need open-ended evaluation or have specific benchmark needs.Votes: 0GitHub stars: 9
- Mmmu A Massive Multi Discipline Multimodal Understanding And Reasoning Benchmark Arxiv 2311 16502v3Use this skill when you need a benchmark testing expert-level multimodal understanding across 30+ subjects using college exam questions. Avoid it when you need basic visual QA evaluation rather than expert-level assessment.Votes: 0GitHub stars: 9
- Molmo And Pixmo Open Weights And Open Data For State Of The Art Multimodal Models Arxiv 2409 17146v1Use this skill when you want fully open-source VLM training data (PixMo) that powers a state-of-the-art model without proprietary data dependencies. Avoid it when you are fine using proprietary data sources.Votes: 0GitHub stars: 9
- Msr Vtt A Large Video Description Dataset For Bridging Video And Language Arxiv Msrvtt 2016Use this skill when you need a video description dataset with 10K web clips and 200K sentences for training and evaluating video-language models. Avoid it when you have sufficient video captioning data.Votes: 0GitHub stars: 9
- Multimodal C4 An Open Billion Scale Corpus Of Images Interleaved With Text Arxiv 2304 06939v2Use this skill when you want to augment C4 text corpus with interleaved images placed at relevant positions for multimodal pretraining. Avoid it when you only need paired image-text data or cannot process the billion-scale corpus.Votes: 0GitHub stars: 9
- Nlvr2 A Visual Reasoning Benchmark For Natural Language Arxiv 1811 00491v2Use this skill when you need a visual reasoning benchmark where the model must determine if a statement is true for a pair of images. Avoid it when single-image reasoning evaluation is sufficient.Votes: 0GitHub stars: 9
- No Robots A Dataset Of Personally Written Instructions Arxiv 2312 15233v1Use this skill when you want a dataset of 10K personally written instruction-response pairs created entirely by skilled humans, not AI. Avoid it when AI-generated instruction data is acceptable for your needs.Votes: 0GitHub stars: 9
- Nocaps Novel Object Captioning At Scale Arxiv 1812 08658v2Use this skill when you need a captioning benchmark testing generalization to novel objects not seen during training. Avoid it when you only evaluate on in-domain captioning.Votes: 0GitHub stars: 9
- Nuscenes A Multimodal Dataset For Autonomous Driving Arxiv 1903 11027v5Use this skill when you need a comprehensive multimodal autonomous driving dataset with cameras, LiDAR, radar, and 3D annotations for driving VLMs. Avoid it when driving data is not relevant to your work.Votes: 0GitHub stars: 9
- Obelics An Open Web Scale Filtered Dataset Of Interleaved Image Text Documents Arxiv 2306 16527v2Use this skill when you want to build an open web-scale dataset of interleaved image-text documents for multimodal pretraining. Avoid it when you only need paired image-caption data without document context.Votes: 0GitHub stars: 9
- Objects365 A Large Scale High Quality Dataset For Object Detection Arxiv 1908 07540v1Use this skill when you need a large-scale detection dataset with 365 categories for training open-vocabulary detection models used in VLM pipelines. Avoid it when you only need the 80 COCO categories or a smaller detection dataset.Votes: 0GitHub stars: 9
- Ocr Vqa Visual Question Answering By Reading Text In Images Crossref 10 1109 Icdar 2019Use this skill when you need a VQA dataset about book covers requiring OCR to read titles, authors, and other text. Avoid it when you need general VQA or document VQA rather than book cover understanding.Votes: 0GitHub stars: 9
- Ok Vqa A Visual Question Answering Benchmark Requiring External Knowledge Arxiv 1906 00067v2Use this skill when you need a VQA dataset where questions require external knowledge sources like Wikipedia to answer. Avoid it when your VQA task does not require external knowledge retrieval.Votes: 0GitHub stars: 9
- Open X Embodiment Robotic Learning Datasets And Rt X Models Arxiv 2310 08864v2Use this skill when you want to aggregate robot learning data across different embodiments and labs for training generalist robot policies. Avoid it when you are not working with robotics or have data from a single robot type.Votes: 0GitHub stars: 9
- Openimages V7 Extended Dataset With Bounding Boxes And Visual Relationships Crossref Openimages 2020Use this skill when you need a massive multi-label detection dataset with 9M images, 600 categories, and visual relationship annotations. Avoid it when COCO or Objects365 provide sufficient detection coverage.Votes: 0GitHub stars: 9