All authors

Claude Skills by ADu2021
github.com/ADu20211,228 skills0 installs1,664 views
- Graph Optimization Test Time ComputeOptimize test-time computation through graph-based collaborative architecture where nodes represent models, edges represent information flow, and topology itself is optimizable via reinforcement learning to discover ideal model assignments and configurations.Votes: 0GitHub stars: 6
- Green Vla RoboticsTrain robot controllers via five-stage curriculum progressing from base vision-language models to embodiment-specific RL-refined policies. Unified action space enables cross-embodiment transfer with minimal performance loss.Votes: 0GitHub stars: 6
- Group Rank Reranking RlImprove ranking quality via groupwise reranking with RL—process document groups jointly for within-group comparisons using GRPO with composite rewards (recall, ranking metrics, distribution alignment).Votes: 0GitHub stars: 6
- Grove Moe Adjugate ExpertsEnables efficient MoE architectures through heterogeneous expert sizing and dynamic activation mechanisms that adjust parameter count based on input complexity.Votes: 0GitHub stars: 6
- Growing Transformers Modular ExpansionBuild language models layer-by-layer on frozen embeddings, training new Transformer blocks sequentially while keeping lower layers frozen. Achieves 5% improvement over monolithic baselines on MMLU while fitting 740M trainable parameters per stage on single GPUs, enabling resource-efficient incremental scaling to 2.3B parameters.Votes: 0GitHub stars: 6
- Grpo Ma Multi Answer Cot TrainingStabilize and accelerate chain-of-thought RL training by sampling multiple answers per generated thought. GRPO-MA reduces gradient noise and improves convergence across math, code, vision, and manipulation tasks while cutting computational cost versus naive thought scaling—critical for training reasoning models without explicit value networks.Votes: 0GitHub stars: 6
- Gtr Turbo VlmEliminate expensive external teacher dependencies in VLM RL training via merged-checkpoint teachers. Uses TIES merging of historical RL checkpoints to create free, stable teacher models for step-level guidance—matching external teacher performance while reducing training time 50% and computational costs 60%.Votes: 0GitHub stars: 6
- Guardians Of Hair Soft BoundariesRecover fine details at soft boundaries (hair, fur) through depth refinement networks and view synthesis. Integrate plug-and-play with existing depth models via adaptive combination across monocular, stereo, and novel view tasks.Votes: 0GitHub stars: 6
- Gui 360 Desktop Agent DatasetEnable training and evaluation of desktop computer-using agents through 1.2M action steps across diverse Windows applications, covering GUI grounding, screen parsing, and action prediction with hybrid GUI+API action space reflecting modern agent designs.Votes: 0GitHub stars: 6
- Gui Actor GroundingEnable GUI agents to ground actions without generating pixel coordinates by using attention-based patch-level alignment and a verifier for selecting optimal action regions from candidates.Votes: 0GitHub stars: 6
- Gui Test Time ScalingImprove GUI agent planning and action grounding through test-time scaling and reinforcement learning. Sample and evaluate multiple action candidates, then use RL to precisely target visual interface elements.Votes: 0GitHub stars: 6
- Guidelines To Prompt Large Language Models For CodImplement techniques from Guidelines to Prompt Large Language Models for Code Generation: An Empirical Characterization. Large Language Models (LLMs) are nowadays extensively used for various types of software engineering tasks, primarily code generationVotes: 0GitHub stars: 6
- H Net Dynamic Chunking Hierarchical SequenceEliminate fixed tokenization by learning data-dependent segmentation jointly with the model through dynamic chunking, matching BPE-tokenized Transformers at equivalent compute while showing improved robustness and better downstream task performance without vocabulary constraints.Votes: 0GitHub stars: 6
- Hardtests Code VerificationGenerate comprehensive test cases for code problems that reliably detect wrong solutions through LLM-based edge case synthesis and test quality ranking.Votes: 0GitHub stars: 6
- Hardware Agnostic Reranker EvaluationEvaluate LLM-based document rerankers using hardware-agnostic FLOPs metrics instead of latency, enabling fair comparison of ranking quality per unit of computation across different models and deployment scenarios.Votes: 0GitHub stars: 6
- Harmonyguard Safety Utility AgentsMulti-agent framework that balances safety compliance with task completion through adaptive policy extraction and dual-objective optimization. Achieves 38% improvement in policy compliance while maintaining 20% higher task completion.Votes: 0GitHub stars: 6
- Hcapo Hindsight Credit AssignmentCompute step-level credit assignments via hindsight generative verification: condition the LLM on successful outcomes to compute importance ratios that filter credit by causal relevance. Synergizes macro stability with micro precision.Votes: 0GitHub stars: 6
- Hermes Hierarchical Video MemoryUse KV cache as hierarchical memory for real-time video stream understanding with minimal GPU overhead, achieving 10x faster response times compared to standard methods. Use when processing continuous video streams where latency and memory efficiency are critical.Votes: 0GitHub stars: 6
- Heteroscale AutoscalingScale disaggregated LLM inference (prefill-decode) via topology-aware scheduling and metric-driven policies, achieving 26.6% GPU utilization improvement and conserving hundreds of thousands GPU-hours dailyVotes: 0GitHub stars: 6
- Hierarchical Lvm ReasoningReason about long-horizon dynamics by disentangling structure and motion in video VAE latents. Learn continuous latent motion chains that preserve temporal coherence while predicting terminal keyframes, enabling efficient reasoning about multi-step scenarios.Votes: 0GitHub stars: 6
- High Entropy Minority Tokens RlOptimize only high-entropy tokens during RL training to achieve better reasoning performance with 80% fewer gradient updates.Votes: 0GitHub stars: 6
- Higher Order Linear Attention MechanismEnable data-dependent higher-order interactions in attention using prefix-sufficient statistics that maintain linear time and constant state, replacing quadratic dot-product attention while preserving expressivity through compact matrix operations.Votes: 0GitHub stars: 6
- Himap TravelSolves long-horizon planning problems with global constraints by decoupling planning into strategic (resource allocation) and tactical (execution) levels. Prevents constraint drift through synchronized state tracking and cooperative bargaining.Votes: 0GitHub stars: 6
- Hop Skip Overthink DiagnosisNovel error categorization framework examining failures across hops (diversity), coverage, and overthinking. Combines human annotation with automated metrics to diagnose why reasoning models hallucinate on multi-step tasks.Votes: 0GitHub stars: 6
- Hpsv3 Human Preference Score EvaluationA VLM-based preference scoring system trained on 1.17M annotated comparisons to evaluate text-to-image generation quality at scale. Uses uncertainty-aware ranking loss for fine-grained assessment across diverse images and supports iterative quality improvement through chain-of-human-preference sampling.Votes: 0GitHub stars: 6
- Humanomni Multimodal ReasoningImprove multimodal reasoning by requiring explicit context understanding before reasoning. Use specialized reward mechanisms and context-aware training to prevent information-skipping shortcuts.Votes: 0GitHub stars: 6
- Hybrid Linear AttentionDesign hybrid architectures mixing linear and full attention at optimal ratios. Superior standalone linear models don't necessarily excel in hybrids; recall improves significantly with more full attention layers at ratios below 3:1, enabling efficient long-sequence processing.Votes: 0GitHub stars: 6
- Hybrid Reinforcement LearningCombine sparse verifier rewards with dense reward model scores using stratified normalization to overcome limitations of either approach alone.Votes: 0GitHub stars: 6
- Hyperagents Self ImprovementEnable AI systems to recursively improve themselves by making the meta-level modification procedure itself editable, achieving open-ended capability growth.Votes: 0GitHub stars: 6
- Hypergraph Memory RagBuild hypergraph-structured memory systems for multi-step RAG that capture high-order relationships between facts, enabling stronger reasoning across long contexts. Use when combining multiple retrieved documents in complex reasoning chains that require understanding connections between pieces of information.Votes: 0GitHub stars: 6
- I Grpo Self Feedback ReasoningImprove LLM reasoning through iterative refinement where the model refines its best previous attempts. Two-stage training: exploratory draft generation, then conditioned refinement using GRPO. Dynamic conditioning signals evolve with policy, enabling state-of-the-art math reasoning on AIME (85.62%).Votes: 0GitHub stars: 6
- Ieap Image EditingEnable robust image editing by decomposing free-form instructions into sequential atomic operations executed through a neural program interpreter.Votes: 0GitHub stars: 6
- Ifdecorator Instruction Following RlEnhances RLVR through cooperative-adversarial flywheel, intent verification, and trap instructions. Detects reward hacking and improves training efficiency, achieving 87.43% on IFEval.Votes: 0GitHub stars: 6
- Illusion Of ThinkingEvaluate reasoning model capabilities by analyzing three complexity-dependent behavioral regimes and identifying fundamental limitations in symbolic manipulation rather than computational budgets.Votes: 0GitHub stars: 6
- Image Super Resolution AgentsUpscale any degraded image to 4K using an agentic framework that analyzes image quality, selects appropriate restoration tools, and iteratively improves results through reasoning and reflection.Votes: 0GitHub stars: 6
- Imagine Then Plan World ModelImagine-then-Plan: Agent Learning from Adaptive Lookahead with World Models. From arXiv:2601.08955Votes: 0GitHub stars: 6
- Indexcache Sparse Attention AccelerationAccelerate sparse attention by reusing token selection indices across layers. Partition layers into full indexer (F) and shared (S) types using greedy search or multi-layer distillation to eliminate 75% of indexer computation.Votes: 0GitHub stars: 6
- Industrial Defect Multimodal DatasetIntroduce IMDD-1M dataset with 1M aligned image-text pairs spanning 60+ material categories and 400+ defect types. Train diffusion-based vision-language models requiring 5% task-specific data vs. dedicated expert models for manufacturing quality control.Votes: 0GitHub stars: 6
- Inference Time Scaling Of Verification Self EvolviImplement techniques from Inference-Time Scaling of Verification: Self-Evolving Deep Research Agents via Test-Time Rubric-Guided Verification. While the majority of existing efforts focus on enhancing policy capabilities via post-training, we propose an alternative paradigm: self-evolving the agent's ability by iteratively verifying the policy model's outputs, guided by meticulously crafted rubricsVotes: 0GitHub stars: 6
- Infialign Data SelectionCombines SFT and DPO with robust data selection pipeline using multidimensional quality metrics. Achieves DeepSeek-R1 performance with 12% training data, enabling efficient reasoning model alignment.Votes: 0GitHub stars: 6
- Infinidepth Arbitrary Resolution And Fine GrainedResearch contribution advancing agent and reasoning capabilities through novel approaches to model development, training, and evaluation.Votes: 0GitHub stars: 6
- Infinitevl Linear Sparse AttentionMerge sliding window and linear attention (Gated DeltaNet) for unlimited VLM inputs with 3.6× speedup. InfiniteVL handles video understanding at 24 FPS with constant memory—ideal when context length must scale without quadratic overhead.Votes: 0GitHub stars: 6
- Infllm V2 Dense Sparse Switchable AttentionImplement dense-sparse switchable attention enabling LLMs to scale from short to long sequences with 4× speedup and 98-99.7% performance retention, requiring no extra parameters by reusing pretrained attention weights through trainable sparse pattern selection.Votes: 0GitHub stars: 6
- Info Driven Policy Optimization AgentsOptimize multi-turn agent policies by measuring turn-level information gain via counterfactual reasoning. Provide dense reward signals identifying which clarifying questions and observations improve the agent's decision distribution, then adaptively blend information rewards with outcome rewards.Votes: 0GitHub stars: 6
- Inplace Feedback Multi Turn ReasoningEnable more precise LLM error correction by having users directly edit the model's previous response, conditioning the next response on this corrected version. This approach reduces token overhead by 79% compared to traditional separate-feedback methods while fixing more errors in complex reasoning tasks.Votes: 0GitHub stars: 6
- Insight O3 MultimodalEnable VLMs to perform generalized visual search—locating relational, fuzzy, and conceptual regions from free-form language descriptions. Introduces O3-Bench benchmark with high-density composite charts/maps, uses RL-trained vSearcher for spatial localization, improving frontier models (GPT-5-mini 39%→61.5%) without architecture changes.Votes: 0GitHub stars: 6
- Interactive Video GenerationConvert pre-trained latent video diffusion into real-time autoregressive generators using adversarial post-training, achieving 24fps streaming on single H100.Votes: 0GitHub stars: 6
- Internvl3 5 Multimodal Cascade RlEnhance multimodal models through cascade RL for reasoning improvement and visual resolution routing for inference efficiency, achieving 16% reasoning gains and 4.05x speedup.Votes: 0GitHub stars: 6
- Inverse Llava Text To VisionMap text embeddings into visual representation space for multimodal fusion, eliminating expensive image-text alignment pre-training while improving reasoning-heavy tasks by up to 27.2%.Votes: 0GitHub stars: 6
- Iot Mcp Llm Device InteractionConnect LLM agents to IoT sensors and microcontrollers through MCP standardization. Use to build monitoring systems and smart home automation where LLMs reason over real-world sensor data.Votes: 0GitHub stars: 6