All authors

Claude Skills by wenmin-wu
github.com/wenmin-wu535 skills0 installs688 views
- Multi Seed Fold AveragingTrains multiple models per CV fold with different random seeds for augmentation, then averages their predictions to reduce variance from stochastic data generation.Votes: 0GitHub stars: 61
- Multi Source Candidate FusionFuse recommendation candidates from user history, multiple co-visitation matrices, and global popularity in a priority-ordered cascadeVotes: 0GitHub stars: 61
- Multiclass To Binary CollapseTrains on a finer-grained multiclass target (subtypes), then collapses non-baseline classes into a single positive class for binary submission.Votes: 0GitHub stars: 61
- Nelder Mead Threshold OptimizationUse scipy Nelder-Mead simplex to optimize regression-to-ordinal thresholds maximizing quadratic weighted kappa on OOF predictionsVotes: 0GitHub stars: 61
- Next Click Time DeltaComputes seconds until the next event within a group using diff().shift(-1) on sorted timestamps, capturing user behavior velocity.Votes: 0GitHub stars: 61
- Ngram Composite FeaturesCreates bi-gram and tri-gram composite categorical features by concatenating top categorical columns, then target-encodes the composites. Captures interaction effects that tree models may miss.Votes: 0GitHub stars: 61
- Null Importance Feature SelectionScores features by comparing actual importances against a null distribution from shuffled targets, removing features that cannot beat random noise.Votes: 0GitHub stars: 61
- Oof Meta FeaturesGenerates out-of-fold predictions from auxiliary models and uses them as input features for the final model.Votes: 0GitHub stars: 61
- Opponent Avoidance ScoringScore candidate movement directions by average distance to nearby opponents and pick the safest path for ball-carrying agents in game AIVotes: 0GitHub stars: 61
- Optuna Lgbm TuningUses Optuna with TPE sampler for Bayesian hyperparameter optimization of LightGBM, searching key params like num_leaves, depth, and learning rate.Votes: 0GitHub stars: 61
- Outlier Aware Two Stage BlendingWhen a regression target has a long discrete tail (e.g. ~1% of rows pinned at -33.22 in Elo), train one regressor on the *non-outlier* subset, a separate binary classifier for the outlier flag, and splice the predictions — replace the top-K most-confident outlier predictions in the regressor's output with the outlier value, where K is calibrated on validationVotes: 0GitHub stars: 61
- Outlier Rate Label EncodingEncode a categorical column by replacing each category with the per-category outlier rate (mean of a binary outlier flag), out-of-fold to avoid leakage — a target-aware encoding tuned to long-tail / sentinel-target problems where a binary classifier signal is more useful than the raw regression meanVotes: 0GitHub stars: 61
- Pairwise Te Logit StackingGenerates all C(n,2) pairwise feature combinations, target-encodes each pair with cuML TargetEncoder, then applies logit polynomial expansion (z, z^2, z^3) for stacking with cuML LogisticRegression.Votes: 0GitHub stars: 61
- Pearson Correlation LossUses negative row-wise Pearson correlation as a differentiable loss function for multi-output regression, directly optimizing the competition metric.Votes: 0GitHub stars: 61
- Per Feature Bias CorrectionPost-processing correction for multi-output regression — scale each output by its train-derived mean ratio to fix systematic per-feature biasVotes: 0GitHub stars: 61
- Per Fold Threshold Voting EnsembleBinarizes each CV fold's predictions using its own optimized threshold, then majority-votes across folds instead of averaging raw probabilities.Votes: 0GitHub stars: 61
- Per Partition Variance FilteringApply VarianceThreshold within each data partition on combined train+test to select informative features per subgroupVotes: 0GitHub stars: 61
- Per Target Nan Mask TrainingTrains independent models per target by masking NaN labels, enabling multi-output regression on datasets where each target has different coverage.Votes: 0GitHub stars: 61
- Per Type Model TrainingTrains separate models for each discrete category (e.g., molecule type, product class) to capture type-specific patterns.Votes: 0GitHub stars: 61
- Personnel Count ParsingParse structured text fields like '1 RB, 2 TE, 2 WR' into separate numeric columns per categoryVotes: 0GitHub stars: 61
- Phase Based Strategy CyclingDivide a game into repeating phases (attack, mine, spawn) with turn-modular gating so the agent cycles between aggressive and economic behaviorVotes: 0GitHub stars: 61
- Play Direction NormalizationMirror spatial coordinates and angles so all plays face the same direction — removes left/right asymmetry from sports and spatial dataVotes: 0GitHub stars: 61
- Polynomial Interaction FeaturesGenerates polynomial powers and interaction terms from selected numeric features to capture nonlinear relationships with the target.Votes: 0GitHub stars: 61
- Popularity Fallback RecommendationFills unfilled recommendation slots with globally popular recent items to handle cold-start users and short lists.Votes: 0GitHub stars: 61
- Ppo Gym Wrapper Kaggle EnvWrap a Kaggle competitive game environment as an OpenAI Gym env with continuous action space for training PPO agents via stable-baselines3Votes: 0GitHub stars: 61
- Predicted Class Mass ReweightingPost-hoc rescales ensemble probabilities by the inverse of each class's estimated total mass across the test set, correcting for class imbalance in predictions.Votes: 0GitHub stars: 61
- Prior Rebalancing OversamplingRebalances training data by oversampling the majority class to match a known test-set class prior, reducing prediction miscalibration.Votes: 0GitHub stars: 61
- Pseudo LabelingAugments training data with high-confidence test predictions as pseudo labels, retrains the model, and keeps the result only if OOF AUC improves. A semi-supervised technique for tabular competitions.Votes: 0GitHub stars: 61
- Rank Averaging EnsembleEnsembles multiple model predictions by converting to ranks, averaging, and normalizing back to [0,1].Votes: 0GitHub stars: 61
- Rank Calibrated BlendingBlends predictions from multiple models by converting to ranks, weighting, and calibrating back to probabilities via rank-group means from a reference model. Ensures monotonic calibrated output.Votes: 0GitHub stars: 61
- Rdkit Molecular DescriptorsComputes all numeric RDKit molecular descriptors from SMILES strings, filtering out NaN, constant, and infinite values to produce a clean feature matrix.Votes: 0GitHub stars: 61
- Recency Weighted Candidate GenerationGenerates recommendation candidates by ranking a customer's purchase history by frequency and recency within a recent window.Votes: 0GitHub stars: 61
- Rectangular Flight Plan EncodingEncode closed rectangular patrol routes as compact direction-distance strings for fleet pathfinding on toroidal game gridsVotes: 0GitHub stars: 61
- Recursive Feature EliminationUses RFE with a tree estimator to iteratively remove least important features, selecting an optimal compact feature set.Votes: 0GitHub stars: 61
- Regression To Cdf SmoothingConvert a scalar regression prediction into a smoothed CDF over discrete bins using a linear ramp instead of a hard stepVotes: 0GitHub stars: 61
- Regression To Ordinal ThresholdingConverts regression predictions to ordinal classes by optimizing bin thresholds to maximize Quadratic Weighted Kappa.Votes: 0GitHub stars: 61
- Regularized Qda ClassifierUse QuadraticDiscriminantAnalysis with regularization for binary classification on data with Gaussian cluster structureVotes: 0GitHub stars: 61
- Relative Deviation FeaturesComputes differences and ratios between group-level aggregates and raw values to capture how each sample deviates from its group.Votes: 0GitHub stars: 61
- Ridge Xgb StackingTwo-stage stacking where Ridge regression on OHE+scaled features produces OOF predictions fed as an extra feature to XGBoost, letting the tree model correct non-linear residuals on top of captured linear patterns.Votes: 0GitHub stars: 61
- Rmsle Keras MetricCustom Keras RMSLE metric using K.log with K.clip to safely evaluate price and count regression during trainingVotes: 0GitHub stars: 61
- Row Aggregate FeaturesEngineers row-wise statistical features (sum, mean, std, skew, kurtosis, median, min, max) across all numeric columns per sample.Votes: 0GitHub stars: 61
- Row Wise Target NormalizationNormalizes each sample's multi-output target vector to zero mean and unit variance, removing per-sample scale differences before training.Votes: 0GitHub stars: 61
- Season Phase LabelingMap calendar dates to categorical season phases (offseason, preseason, regular, postseason) using np.select with boundary date conditionsVotes: 0GitHub stars: 61
- Simulated Annealing Multi OperatorSimulated annealing with diverse move operators (translate, rotate, swap, Levy flight, squeeze) and adaptive reheating on stagnation for combinatorial optimizationVotes: 0GitHub stars: 61
- Smiles Randomization AugmentationAugments molecular datasets by generating multiple randomized SMILES strings for the same molecule, exploiting SMILES non-uniqueness to multiply training samples.Votes: 0GitHub stars: 61
- Sparse Dense Hstack LgbmTrain LightGBM directly on a scipy.sparse.hstack of TF-IDF text vectors and dense tabular columns, passing feature_name and categorical_feature so native categorical handling survives the sparse blockVotes: 0GitHub stars: 61
- Spatial Distance AggregationCompute min/max/mean/std of Euclidean distances from all entities to a key point, then aggregate per group for spatial feature engineeringVotes: 0GitHub stars: 61
- Squeeze Compact Local SearchThree-stage packing refinement — uniform squeeze toward centroid, greedy compaction per object, then multi-directional local search — to tighten solutions after metaheuristic optimizationVotes: 0GitHub stars: 61
- Streaming Prediction ApiOnline inference pattern that processes test batches sequentially, updating feature dictionaries incrementally for time-series prediction APIs.Votes: 0GitHub stars: 61
- Strtree Spatial Index CollisionUse Shapely STRtree spatial index for O(n log n) polygon overlap detection instead of brute-force O(n^2) pairwise checksVotes: 0GitHub stars: 61