All authors

Claude Skills by wenmin-wu
github.com/wenmin-wu535 skills0 installs688 views
- Greedy Word Reinsert SearchGreedy local search that removes one element from a fixed position and re-inserts it at every possible index, keeping the best improvement per roundVotes: 0GitHub stars: 61
- Inverse Task Prompt TemplateStructured prompt template for recovering the instruction that transformed one text into another, with labeled original/rewritten fields and explicit task framingVotes: 0GitHub stars: 61
- Iterative Candidate RerankingNarrows a large candidate pool through multiple LLM voting rounds, each presenting a sliding window of candidates plus the current best pick.Votes: 0GitHub stars: 61
- Iterative Pairwise Keyword ExtractionIteratively prompts an LLM over document pairs to extract and deduplicate keywords, building a comprehensive term set from multiple perspectives.Votes: 0GitHub stars: 61
- Kv Cache Prefix ScoringShares KV cache from a common prefix (context+question) across multiple answer suffixes for efficient multi-choice scoring.Votes: 0GitHub stars: 61
- Last Token Logit Binary ScoringScore a binary classification prompt by reading the logits of the True/False (or Yes/No) token IDs at the final position and softmaxing only those two values, skipping generation entirely for a 10-50x speedup over decodingVotes: 0GitHub stars: 61
- Last Token Pooling EmbeddingExtracts dense sentence embeddings from decoder-only LLMs by pooling the last non-padding token's hidden state.Votes: 0GitHub stars: 61
- Lora Sequence ClassificationLoad a pretrained LLM with LoRA adapter via PEFT for memory-efficient fine-tuned sequence classificationVotes: 0GitHub stars: 61
- Low Meaning Input SubstitutionReplace the actual input text with a generic low-meaning passage to prevent the LLM from fixating on content specifics, forcing it to focus on stylistic and structural transformation cuesVotes: 0GitHub stars: 61
- Multi Gpu Process Isolated VllmRun independent vLLM workers on each GPU by spawning one mp.Process per device and setting CUDA_VISIBLE_DEVICES inside the child before vLLM is imported, sidestepping vLLM's single-instance-per-process limitationVotes: 0GitHub stars: 61
- Multiple Choice Logits ProcessorConstrains LLM generation to a fixed set of valid choice tokens using a logits processor for structured single-token output.Votes: 0GitHub stars: 61
- Perplexity Prompt RankingRank candidate prompts by computing LLM perplexity of the full conversation conditioned on each prompt, selecting the lowest-perplexity candidate as the best matchVotes: 0GitHub stars: 61
- Prompt Variant EnsembleGenerate multiple LLM responses using diverse system prompt variants to increase reasoning diversity for self-consistency votingVotes: 0GitHub stars: 61
- Pseudo Metadata Score InjectionAppend fake metadata tags like [Score: 8.7] or [plagiarism_odds_pct: 95.2] to anchor an LLM judge's numeric outputVotes: 0GitHub stars: 61
- Score Variance Baseline MixinMix a small fraction of plain baseline responses into an adversarial submission to preserve cross-row score varianceVotes: 0GitHub stars: 61
- Self Consistency Majority VoteAggregate multiple LLM reasoning attempts via majority voting with random jitter tiebreaking and validity filteringVotes: 0GitHub stars: 61
- Sentence Truncation FallbackTruncate LLM output to exactly N sentences and fall back to a known-good baseline string when output is empty or too shortVotes: 0GitHub stars: 61
- Sliding Window Permutation SearchLocal search that slides a window of size p across a word sequence, brute-forcing all permutations within each window to minimize an objective like LLM perplexityVotes: 0GitHub stars: 61
- Spiral Patrol ExplorationAgents patrol in expanding spiral patterns using rotating direction sequences with increasing radius for systematic grid explorationVotes: 0GitHub stars: 61
- Stopword Priority SortingInitialize text ordering by placing stopwords first then content words, producing low-perplexity starting points for combinatorial search over word permutationsVotes: 0GitHub stars: 61
- Structured Output SanitizationParses LLM-generated markup (SVG, HTML, XML) with lxml, strips disallowed elements and attributes via an allowlist, and validates structural constraints like path data.Votes: 0GitHub stars: 61
- Svg Constrained GenerationPrompts an LLM to generate valid SVG by embedding an explicit element/attribute allowlist and a one-shot example, then extracts the last valid SVG block from output.Votes: 0GitHub stars: 61
- Swarm Tactic DiversityAssign each new agent a different directional rotation pattern from a set of permutations to ensure swarm coverage diversity across the mapVotes: 0GitHub stars: 61
- Synthetic Data AugmentationGenerates additional training examples using a stronger LLM (e.g., GPT-3.5) to augment small labeled datasets.Votes: 0GitHub stars: 61
- Test Time Train Pseudo Label ExpansionConvert a test.csv with paired positive/negative example columns into a labeled training set at inference time, using the OTHER example as the in-prompt demonstration so the model never sees its own target as a few-shot exemplarVotes: 0GitHub stars: 61
- Tfidf Chunked RetrievalScalable TF-IDF retrieval over large document corpora using frozen vocabulary and chunked top-k merging.Votes: 0GitHub stars: 61
- Threaded Multi Gpu InferenceRun multiple LLM inference jobs in parallel using Python threads, each pinned to a separate GPU with staggered startsVotes: 0GitHub stars: 61
- Turn Based Prompt AccumulationRebuild multi-turn conversation context by interleaving user/assistant turns with chat template tokens into a single prompt each callVotes: 0GitHub stars: 61
- Two Pass Retrieval RefinementRefines retrieval by running two passes: initial embedding retrieval to get candidates, then LLM-generated text concatenated with the query for a second retrieval pass.Votes: 0GitHub stars: 61
- Vllm Lora Adapter InferenceServe a quantized base LLM with a hot-swappable LoRA adapter under vLLM, enabling prefix caching and tensor parallelism so a single fine-tuned adapter runs at production throughput without merging weightsVotes: 0GitHub stars: 61
- Weighted Attack Strategy SamplingSample from a pool of adversarial prompt strategies with per-strategy probability weights to hedge across judge modelsVotes: 0GitHub stars: 61
- Wikipedia Rag RetrievalDense retrieval over a FAISS-indexed Wikipedia corpus to provide grounding context for LLM question answering.Votes: 0GitHub stars: 61
- 8bit Optimizer Embedding OverrideUses bitsandbytes 8-bit AdamW to halve optimizer memory, with a 32-bit override for embedding weights to prevent instability.Votes: 0GitHub stars: 61
- Adafactor Label Smoothing Seq2seqUse Adafactor optimizer with label smoothing for seq2seq fine-tuning — memory-efficient and regularizes overconfident predictionsVotes: 0GitHub stars: 61
- Adjacent Span MergingMerges nearby predicted NER spans of the same class within a word-distance threshold into single coherent segments.Votes: 0GitHub stars: 61
- Anchor Grouped ValidationSplit validation by unique anchor/query entities so no anchor appears in both train and val, preventing data leakage in pairwise matching tasksVotes: 0GitHub stars: 61
- Attention Head PoolingLearns attention weights over token positions to compute a weighted average of hidden states for sequence representation.Votes: 0GitHub stars: 61
- Auxiliary Target MultitaskTrain main target alongside auxiliary sub-type targets as multiple output heads to regularize and improve generalizationVotes: 0GitHub stars: 61
- Averaged Meta EmbeddingElement-wise average of multiple pretrained embedding matrices as a parameter-free meta-embeddingVotes: 0GitHub stars: 61
- Best Prob Fallback MatchingWhen no candidate passes the threshold for a query, fall back to the single highest-scoring match to guarantee at least one prediction per queryVotes: 0GitHub stars: 61
- Bias Aware Auc MetricCombine subgroup AUC, BPSN AUC, and BNSP AUC across identity groups via power-mean weighting with overall AUC for fairness evaluationVotes: 0GitHub stars: 61
- Bidirectional Translation AugmentationDouble training data by adding reverse-direction translation pairs with task prefix promptsVotes: 0GitHub stars: 61
- Bio Tag Span ReconstructionReconstructs named entity spans from BIO token-level tags, handling B/I/O transitions and sentence boundaries.Votes: 0GitHub stars: 61
- Bio Tagging Sliding WindowSplits long documents into overlapping fixed-length windows with BIO NER tags for BERT token classification on sequences exceeding max length.Votes: 0GitHub stars: 61
- Bpe Offset Char AlignmentReconstruct character-level offsets for BPE tokens by decoding each token individually and accumulating lengths for precise span mappingVotes: 0GitHub stars: 61
- Bucket Batching Seq2seqGroup variable-length sequences into length-sorted buckets before batching to minimize padding waste during seq2seq inferenceVotes: 0GitHub stars: 61
- Category Weighted Score FusionConverts multi-label binary flags into a continuous regression target by applying hand-tuned per-category multipliers, then averaging across categories.Votes: 0GitHub stars: 61
- Char Prob Weighted BlendBlend character-level probability arrays from multiple models with OOF-tuned weights before thresholdingVotes: 0GitHub stars: 61
- Checkpoint Ensemble ExponentialAverage predictions from each epoch checkpoint with exponentially increasing weights (2^epoch), favoring later more-converged snapshotsVotes: 0GitHub stars: 61
- Chrf Bleu Geometric Mean MetricGeometric mean of chrF and BLEU as a balanced composite translation evaluation metricVotes: 0GitHub stars: 61