All authors

Claude Skills by feiyang-k
github.com/feiyang-k258 skills0 installs459 views
- The All Seeing Project Towards Panoptic Visual Recognition And Understanding Of The Open World Arxiv 2308 01907v2Use this skill when you want to create a large-scale dataset with region-level recognition and text descriptions for panoptic visual understanding. Avoid it when you do not need region-level understanding or panoptic annotations.Votes: 0GitHub stars: 9
- Towards Open World Segmentation Of Parts Arxiv 2305 06914v3Use this skill when you need part-level segmentation data for training VLMs to understand object parts and fine-grained visual components. Avoid it when object-level segmentation is sufficient.Votes: 0GitHub stars: 9
- Trivialaugment Tuning Free Yet State Of The Art Data Augmentation Arxiv 2103 10158v3Use this skill when you want the simplest possible automated augmentation: randomly select one transform and one magnitude per image, no tuning needed. Avoid it when you prefer tuned augmentation policies.Votes: 0GitHub stars: 9
- Unnatural Instructions Tuning Language Models With Almost No Human Labor Arxiv 2212 09689v2Use this skill when you want to generate diverse instruction data by prompting an LLM with creative seed constraints for novel instruction types. Avoid it when standard self-instruct provides sufficient diversity.Votes: 0GitHub stars: 9
- Video Chatgpt Towards Detailed Video Understanding Via Large Vision And Language Models Arxiv 2306 05424v2Use this skill when you want to generate video instruction data by prompting GPT-3.5 with video descriptions and annotations for video understanding VLM training. Avoid it when you do not work with video or have sufficient video instruction data.Votes: 0GitHub stars: 9
- Visual Instruction Tuning Arxiv 2304 08485v2Use this skill when you want to generate multimodal instruction-following data by prompting a text-only LLM with image captions and bounding boxes. Avoid it when you have abundant human-annotated instruction data or need the LLM to see actual images during data generation.Votes: 0GitHub stars: 9
- Vokenization Improving Language Understanding With Contextualized Visual Grounded Supervision Arxiv 2010 06775v2Use this skill when you want to augment text training data with visual tokens ('vokens') retrieved from an image database to ground language in vision. Avoid it when you do not need visual grounding of text data.Votes: 0GitHub stars: 9
- Wizardlm Empowering Large Language Models To Follow Complex Instructions Arxiv 2304 12244v2Use this skill when you want to progressively increase instruction complexity through evolution prompts (deepening, widening, adding constraints). Avoid it when you need simple instructions or cannot afford iterative evolution.Votes: 0GitHub stars: 9