All authors

Claude Skills by wenmin-wu
github.com/wenmin-wu535 skills0 installs688 views
- Optuna Ensemble WeightsLearns optimal per-model blending weights via Optuna optimization, supporting negative weights for error cancellation.Votes: 0GitHub stars: 61
- Pairwise Margin Ranking LossTrains a transformer with MarginRankingLoss on text pairs (more/less toxic), learning to rank rather than classify when only pairwise preference labels are available.Votes: 0GitHub stars: 61
- Pairwise Ranking ValidationEvaluates ranking models by computing the fraction of preference pairs where the model correctly scores the preferred item higher.Votes: 0GitHub stars: 61
- Pairwise Softmax Rank AggregationConverts pairwise binary predictions into continuous ranks via temperature-scaled softmax weighted sum over anchor positions.Votes: 0GitHub stars: 61
- Pearson Correlation MetricUse Pearson correlation coefficient as evaluation metric for semantic similarity regression tasks, selecting best checkpoint by correlation rather than lossVotes: 0GitHub stars: 61
- Per Class Span FilteringFilters predicted NER spans using per-class minimum word-count and mean-probability thresholds to reduce false positives.Votes: 0GitHub stars: 61
- Prompt Text ConcatenationPrepends prompt question or full prompt text to input with SEP token for context-aware text evaluation.Votes: 0GitHub stars: 61
- Quantized Probability StorageStores model softmax probabilities as uint8 (0-255) to reduce RAM by 4x during multi-model ensemble inference.Votes: 0GitHub stars: 61
- Readability Metric FeaturesExtracts established readability scores (Flesch, Gunning FOG, ARI, Coleman-Liau) as numeric features from text.Votes: 0GitHub stars: 61
- Regex Hybrid Ner FallbackSupplements transformer NER predictions with regex-based detection for structured entities (email, phone, URL), aligning regex matches back to token indices via subsequence search.Votes: 0GitHub stars: 61
- Sentence Alignment AugmentationSplit multi-sentence parallel pairs into aligned sentence pairs to expand training data for seq2seq modelsVotes: 0GitHub stars: 61
- Sentiment Conditioned Qa SpanPrepend a sentiment token as the query in a QA-style input to condition span extraction on sentiment class without architectural changesVotes: 0GitHub stars: 61
- Short To Long Span PromotionPromotes a predicted short answer span to its enclosing long-answer candidate by matching token boundaries against pre-extracted document structure.Votes: 0GitHub stars: 61
- Simulated Annealing Query OptimizationUses simulated annealing to select the optimal subset of candidate terms for a boolean search query, maximizing a retrieval metric like AP@K.Votes: 0GitHub stars: 61
- Source Balanced Stratified FoldStratifies CV folds by both target label AND data source to prevent source-specific bias in each fold.Votes: 0GitHub stars: 61
- Space Aware Span PostprocessingClean up character-level span predictions by removing isolated space characters at span boundariesVotes: 0GitHub stars: 61
- Spacy Custom Ner Span ExtractionTrain per-class spaCy NER models to extract task-specific spans as custom named entities with compounding batch sizesVotes: 0GitHub stars: 61
- Span Overlap F1 MetricEvaluates NER span predictions using bidirectional word-index overlap (>=50% both ways) to compute micro-F1 over predicted vs ground-truth spans.Votes: 0GitHub stars: 61
- Spatial Dropout EmbeddingDrop entire embedding channels consistently across all timesteps — preserves temporal structure better than element-wise dropoutVotes: 0GitHub stars: 61
- Spearman Correlation CallbackCustom Keras callback that evaluates Spearman correlation on validation data each epoch with optional early stopping.Votes: 0GitHub stars: 61
- Spell Correction PreprocessingApplies domain-aware spelling correction before transformer input to separate spelling errors from content quality.Votes: 0GitHub stars: 61
- Sqrt Mse LossUses square root of MSE as training loss to directly optimize for RMSE evaluation metric alignment.Votes: 0GitHub stars: 61
- Structured Prompt FormattingFormat multi-field tabular data into a structured natural language prompt with labeled sections for encoder or LLM classificationVotes: 0GitHub stars: 61
- Subtoken Labeling StrategyControls whether all subtokens or only the first subtoken of each word receive NER labels during training and inference.Votes: 0GitHub stars: 61
- Taxonomy Context EnrichmentEnrich model input by mapping categorical codes to human-readable taxonomy descriptions and concatenating them as context for transformer modelsVotes: 0GitHub stars: 61
- Test Vocabulary AlignmentFits TF-IDF vectorizer on test set first to extract vocabulary, then retrains on train set using that vocabulary for feature consistency.Votes: 0GitHub stars: 61
- Text Overlap FeaturesComputes n-gram overlap counts/ratios and NER entity overlap between reference and generated text as features.Votes: 0GitHub stars: 61
- Tfidf Ngram ClassifierHigh n-gram TF-IDF (3-5 grams) with sublinear TF feeding into a weighted soft-voting ensemble of traditional ML classifiers.Votes: 0GitHub stars: 61
- Tfidf Pair Difference EncodingEncodes text pairs by computing the absolute difference of their TF-IDF vectors, collapsing a pair into a single fixed-length feature vector.Votes: 0GitHub stars: 61
- Tfidf To Boolean QueryConverts TF-IDF top-k terms into field-scoped boolean OR queries for structured document retrieval from a full-text index.Votes: 0GitHub stars: 61
- Tfidf Translation MemoryTF-IDF similarity retrieval from a translation memory with SequenceMatcher reranking as a fallback or ensemble componentVotes: 0GitHub stars: 61
- Tfidf Weighted Word MatchComputes word overlap ratio between two texts weighted by inverse corpus frequency, giving rare shared words more importance than common ones.Votes: 0GitHub stars: 61
- Token Map Filtered Topk SpansFilters candidate span indices through a token map to skip special tokens, then cross-products top-k start/end indices with length constraints.Votes: 0GitHub stars: 61
- Token To Char Span PredictionMap token-level classifier outputs back to character-level spans via offset mapping, thresholding, and contiguous groupingVotes: 0GitHub stars: 61
- Tpu Distribution StrategyAuto-detects TPU vs CPU/GPU at runtime and wraps model construction in the appropriate TensorFlow distribution strategy with scaled batch size.Votes: 0GitHub stars: 61
- Tpu Multicore InferenceDistributes inference across multiple TPU cores using torch_xla, each core writing a CSV shard, then merges shards via groupby mean.Votes: 0GitHub stars: 61
- Train Short Infer Long SequenceTrains a transformer at shorter sequence length for speed, then runs inference at a longer sequence length to capture more context, exploiting position embedding generalization.Votes: 0GitHub stars: 61
- Transformer Embedding SvrExtracts frozen embeddings from multiple pretrained transformers and trains SVR on the concatenated features.Votes: 0GitHub stars: 61
- Transformer Layer FreezingFreezes transformer embedding and lower encoder layers to reduce memory, speed up training, and stabilize fine-tuning.Votes: 0GitHub stars: 61
- Transformer Lgbm StackingTwo-stage stacking: transformer predictions plus hand-crafted features feed into a LightGBM meta-learner.Votes: 0GitHub stars: 61
- Translation Regex PostprocessingMulti-rule regex pipeline to clean seq2seq translation outputs — deduplicate phrases, fix punctuation, remove artifactsVotes: 0GitHub stars: 61
- Triplet Loss Biencoder FinetuningFine-tunes a bi-encoder with triplet loss using retrieval-mined hard negatives for dense similarity search.Votes: 0GitHub stars: 61
- Two Stage Retrieve RerankTwo-stage pipeline where an unsupervised bi-encoder retrieves KNN candidates and a supervised cross-encoder reranks them with sigmoid thresholdingVotes: 0GitHub stars: 61
- Uniform Stride SamplingSamples N items uniformly by stride from a variable-length list, always preserving the first and last elements, to fit long sequences into a fixed token budget.Votes: 0GitHub stars: 61
- Validation Finetune ContinuationAfter training on the primary split, continues fine-tuning on the validation set to adapt the model to the target distribution before inference.Votes: 0GitHub stars: 61
- Weight Decay Bias ExclusionExclude bias terms and LayerNorm weights from weight decay to prevent regularization from distorting normalization layersVotes: 0GitHub stars: 61
- Weighted Layer PoolingLearns a weighted combination of CLS embeddings across all transformer layers instead of using only the last layer.Votes: 0GitHub stars: 61
- Whoosh Fulltext Search IndexBuilds a Whoosh full-text search index over documents and queries it with boolean operators, field scoping, and proximity matching.Votes: 0GitHub stars: 61
- Word2vec Spell CorrectionUses Word2Vec vocabulary rank as a word frequency proxy for Norvig-style spell correction, avoiding the need for a separate frequency corpus.Votes: 0GitHub stars: 61
- Worker Stratified KfoldStratifies CV folds by annotator/worker ID to prevent annotator style leakage across train and validation splits in crowd-sourced datasets.Votes: 0GitHub stars: 61