All authors

Claude Skills by hiyenwong
github.com/hiyenwong9,934 skills5 installs19,223 views
- Arxiv 2609 04021v1 Fly Eval An Evidence Driven Evaluation Protocol Fo**arXiv ID:** 2609.04021v1 **Authors:** Yalun Wu, Junfeng Fang, Jiawei Wang, Haotian Liu, Qijun Yang, Minghan Yang, Hongcheng Guo, Zhoujun Li, Boyang Wang **URL:** http://arxiv.org/abs/2609.04021v1 **Utility Score:** 1.00Votes: 0GitHub stars: 3
- Arxiv 2609 04022v1 Representational Alignment Yields Generalizable Sa**arXiv ID:** 2609.04022v1 **Authors:** Lingyu Li, Yan Teng, Yingchun Wang, Xia Hu **URL:** http://arxiv.org/abs/2609.04022v1 **Utility Score:** 1.00Votes: 0GitHub stars: 3
- Arxiv 2609 04194v1 Legibility Is Not Interpretability Comparing Judge**arXiv ID:** 2609.04194v1 **Authors:** Kevin Du, Alexander Hoyle, Laura Ruis, Acyr Locatelli **URL:** http://arxiv.org/abs/2609.04194v1 **Utility Score:** 1.00Votes: 0GitHub stars: 3
- Arxiv 2609 08275v1 Beyond Coherence Benchmarking Professional Editing**arXiv ID:** 2609.08275v1 **Authors:** Tianyi Zeng, Junchao Liao, Yujie Wei, Ziying Zhang, Litao Li, Tianyi Wang, Zhichao Wei, Shuyao Xu, Wenwen Qiang, Siyu Zhu, Zhenghao Zhang, Long Qin **URL:** http://arxiv.org/abs/2609.08275v1 **Utility Score:** 1.00Votes: 0GitHub stars: 3
- Arxiv 2609 08765v1 Benchmark Scores Are Pipeline Dependent A Reliabil**arXiv ID:** 2609.08765v1 **Authors:** Aymene Berriche, Cathrine Shalby, Mohannad Alhanahnah, Yazan Boshmaf **URL:** http://arxiv.org/abs/2609.08765v1 **Utility Score:** 1.00Votes: 0GitHub stars: 3
- Arxiv 2609 08790v1 Evidence Grounded Retrieval For Investigation Hunt**arXiv ID:** 2609.08790v1 **Authors:** Akash Prakash, Boubakr Nour, Makan Pourzandi, Chadi Assi, Mourad Debbabi **URL:** http://arxiv.org/abs/2609.08790v1 **Utility Score:** 1.00Votes: 0GitHub stars: 3
- Arxiv 2609 08861v1 Api Benchmark Scores Do Not Reliably Transfer To C**arXiv ID:** 2609.08861v1 **Authors:** Jennifer Wang, Joachim Baumann, Daniel E. Ho, Sanmi Koyejo **URL:** http://arxiv.org/abs/2609.08861v1 **Utility Score:** 1.00Votes: 0GitHub stars: 3
- Arxiv 2609 08947v1 One Cylinder A Benchmark For Graph Based Surrogate**arXiv ID:** 2609.08947v1 **Authors:** Théodore Michel, Antoine Campos, Alban Dujardin, Henry Areiza, Philippe Meliga, Elie Hachem **URL:** http://arxiv.org/abs/2609.08947v1 **Utility Score:** 1.00Votes: 0GitHub stars: 3
- Arxiv 2609 09388v1 Xai Refine An Automated Explanation Knowledge Loop**arXiv ID:** 2609.09388v1 **Authors:** Yang Qiao, Junjie Wu, Deqiang Qiu, James J. Lah, Liang Zhao **URL:** http://arxiv.org/abs/2609.09388v1 **Utility Score:** 1.00Votes: 0GitHub stars: 3
- Arxiv 2609 09685v1 Align Hold Experience Alignment For Real Time Hold**arXiv ID:** 2609.09685v1 **Authors:** Zuhao Zhang, Xu Liu, Kai Wan, Zihao Lu, Li Ma, Shuai Li **URL:** http://arxiv.org/abs/2609.09685v1 **Utility Score:** 1.00Votes: 0GitHub stars: 3
- Arxiv 2609 09754v1 Lexagenthallu A Hierarchical Benchmark For Profili**arXiv ID:** 2609.09754v1 **Authors:** Yujin Zhou, Mingxuan Zheng, Chuxue Cao, Huang Yidan, Jiale Chen, Yike Guo, Sirui Han **URL:** http://arxiv.org/abs/2609.09754v1 **Utility Score:** 1.00Votes: 0GitHub stars: 3
- Arxiv 2609 09793v1 How Fragile Is Safety Alignment At Frontier Scale**arXiv ID:** 2609.09793v1 **Authors:** Yi Shi, Tanyu Chen, Kai Shen **URL:** http://arxiv.org/abs/2609.09793v1 **Utility Score:** 1.00Votes: 0GitHub stars: 3
- Arxiv 2609 09853v1 The Era By Eon Benchmark A Generated Enterprise Es**arXiv ID:** 2609.09853v1 **Authors:** Benjamin Gruenbaum, Doron Porat, Assaf Natanzon, Roy Zavida, Chen Dinachi, Or Itzahary **URL:** http://arxiv.org/abs/2609.09853v1 **Utility Score:** 1.00Votes: 0GitHub stars: 3
- Arxiv 2609 10055v1 Ontologyaligner Ontology Aligned Retrieval And Hie**arXiv ID:** 2609.10055v1 **Authors:** Jie Song, Zhichuan Xu, Ziyu Lu, Meng Xiao, Cheng Bi, Yuxin Zhang, Xin Zheng, Xiaoran Li, Qiongfang Cao, Hao Yang, Bairong Shen **URL:** http://arxiv.org/abs/2609.10055v1 **Utility Score:** 1.00Votes: 0GitHub stars: 3
- Aschern At Semeval2020 Task 11 It Takes Three To Tango Roberta Crf And Transfer Learning**arXiv ID:** 2008.02837 **Authors:** Anton Chernyavskiy, Dmitry Ilvovsky, Preslav Nakov **Published:** 2020-08-06T18:45:25Z **Abstract:** We describe our system for SemEval-2020 Task 11 on Detection of Propaganda Techniques in News Articles. We developed ensemble models using RoBERTa-based neural architectures, additional CRF layers, transfer learning between the two subtasks, and advanced post-processing to handle the multi-label nature of the task, the consistency between nested spans, rep...Votes: 0GitHub stars: 3
- Auditing Llm As Judge Reliability Measurement ValidityTreats LLM-as-judge evaluator-replacement ambiguity as a measurement-validity problem. Judge upgrades are not interchangeable. Stronger judges reduce but don't remove position/verbosity bias. Proposes audit trails including dataset slices, bias probes, and error-dependence estimates. Activation: LLM-as-judge, evaluation reliability, measurement validity, evaluator bias, AI evaluation.Votes: 0GitHub stars: 3
- Automated Alignment ResearchersAutomated Alignment Researchers (AARs) methodology — using LLMs to conduct AI alignment research via weak-to-strong supervision, reward hacking mitigation, and PGR metric scoring with Claude Opus-class models.Votes: 0GitHub stars: 3
- Automatic Programming Of Cellular Automata And Artificial Neural Networks Guided By Philosophy**arXiv ID:** 1905.04232 **Authors:** Patrik Christen, Olivier Del Fabbro **Published:** 2019-05-10T16:00:09Z **Abstract:** Many computer models such as cellular automata and artificial neural networks have been developed and successfully applied. However, in some cases, these models might be restrictive on the possible solutions or their solutions might be difficult to interpret. To overcome this problem, we outline a new approach, the so-called allagmatic method, that automatically programs...Votes: 0GitHub stars: 3
- Bcpnn Native ExplainabilityBCPNN (Bayesian Confidence Propagation Neural Network) native explainability framework. First XAI taxonomy for BCPNN, mapping architectural primitives to attribution, prototype, concept, counterfactual, and mechanistic explanations. Introduces 16 architecture-level explanation primitives (P1–P16) computed from quantities the model already maintains, plus 5 design-time Configuration-as- Explanation primitives (Config-P1 to Config-P5). Addresses EU AI Act compliance for brain-like neural networ...Votes: 0GitHub stars: 3
- Benchx Cancer Ai BenchmarkLarge-scale AI benchmarking methodology for cancer detection models. Evaluates tumor-detection AI across tumor size, location, demographic subgroups, and imaging protocols using 85,355 CT scans and 12 models. Use when: benchmarking medical AI models, evaluating cancer detection systems, assessing subgroup fairness in healthcare AI, analyzing CT scan AI performance, building robust tumor detection pipelines.Votes: 0GitHub stars: 3
- Beyond Attack Success Rate Action Graded Severity Scale For Tool Using Ai AgentsAgentic red-teaming benchmarks report whether an injected agent was compromised as a single bit: the attack succeeded, or it did not. We argue that this binary attack-success rate discards the informa. Based on arXiv:2607.07474.Votes: 0GitHub stars: 3
- Beyond Representational Alignment With Brain GuideDerived from arXiv:2606.11893 - Beyond representational alignment with brain-guided language models for robust reasoningVotes: 0GitHub stars: 3
- Causally Emergent Alignment HypothesisCausal emergence (ΦID) predicts and aligns with RL agent reward trajectories. The Causally Emergent Alignment Hypothesis states that successful RL agents exhibit causal emergence that is predictive of final reward early in training and whose representational dynamics align with reward improvement. Use when analyzing RL agent representations, measuring causal emergence in neural networks, studying AI-biology alignment, or investigating ΦID as a learning metric. arXiv: 2605.06746Votes: 0GitHub stars: 3
- Cmms Edge Cluster BenchmarkContinuous Multi-Mode Scheduling(CMMS)基准测试平台用于边缘集群调度算法公平比较。统一控制器接口、闭环负载驱动、双指标SLO评分(原始SLO vs 稳态SLO),揭示控制器排名的配置依赖性和切换成本。Activation: edge cluster scheduling, heterogeneous scheduling, SLO benchmark, CMMS, RL scheduling, adaptive benchmark, edge-cloud continuum.Votes: 0GitHub stars: 3
- Coarse Feedback Visual AlignmentCoarse feedback for human-aligned visual representations. Use when: studying how supervisory signal granularity affects brain alignment in neural networks, designing brain-aligned vision models with minimal supervision, comparing coarse vs fine-grained training objectives, deriving coarse category labels from pretrained embeddings (PCA-based splits), representational similarity analysis (RSA) of neural/behavioral alignment, building AI systems aligned with human perception, or investigating w...Votes: 0GitHub stars: 3
- Collective Alignment Public InputMethodology for incorporating public input into AI model alignment through large-scale surveys and democratic value aggregation.Votes: 0GitHub stars: 3
- Context Access Divide Agentic Inequality ArchitectureFormalizes the Context Access Divide (CAD) as a dimension of agentic inequality operating at the interaction level. Dynamic Context Retrieval vs Manual Attachment causes combinatorial collapse in task-success probability. Proposes contextuality as complement to Sharp et al.'s framework. Activation: agentic inequality, context access, interaction-level architecture, agent fairness, AI equity.Votes: 0GitHub stars: 3
- Contravariance Theory Strong Alignment MinimalContravariance Theory methodology — formal proof that minimal DNN solutions to hard tasks exhibit strong alignment of privileged axes, with alignment "zipping" up the network hierarchy. Bridges NeuroAI convergent evolution theory and brain-DNN comparison methods.Votes: 0GitHub stars: 3
- Creating Trustworthy Llms Dealing With Hallucinations In Healthcare Ai**arXiv ID:** 2311.01463 **Authors:** Muhammad Aurangzeb Ahmad, Ilker Yaramis, Taposh Dutta Roy **Published:** 2023-09-26T20:52:46Z **Abstract:** Large language models have proliferated across multiple domains in as short period of time. There is however hesitation in the medical and healthcare domain towards their adoption because of issues like factuality, coherence, and hallucinations. Give the high stakes nature of healthcare, many researchers have even cautioned against its usage until t...Votes: 0GitHub stars: 3
- Decoding Encoding Alignment CritiqueCritical analysis framework for brain-model alignment methods demonstrating that decoding-based similarity metrics (RSA, CKA, Procrustes) are insensitive to internal functional organization. Based on arXiv:2605.05907 (Bertram et al., 2026). Use when: (1) evaluating representational similarity analysis validity, (2) comparing neural systems across species or models, (3) designing encoding manifold analyses, (4) applying Gromov-Wasserstein distance for neural population comparison, (5) question...Votes: 0GitHub stars: 3
- Deep Double DescentSkill for AI agent capabilitiesVotes: 0GitHub stars: 3
- Dendritic Gain Load Alignment PrincipleGain-Load-Alignment Principle for Dendritic E/I Networks - Framework for understanding when branch-local shunting helps in neural population readout. Analyzes DendriNet architecture with varying integration rules, morphology, and synaptic allocation. Use when studying dendritic computation, shunting inhibition, population coding, or neural readout optimization.Votes: 0GitHub stars: 3
- Diceextended A Robust Approach To Counterfactual Explanations In Machine Learning**arXiv ID:** 2504.19027 **Authors:** Volkan Bakir, Polat Goktas, Sureyya Akyuz **Published:** 2025-04-26T21:22:44Z **Abstract:** Explainable artificial intelligence (XAI) has become increasingly important in decision-critical domains such as healthcare, finance, and law. Counterfactual (CF) explanations, a key approach in XAI, provide users with actionable insights by suggesting minimal modifications to input features that lead to different model outcomes. Despite significant advancements, e...Votes: 0GitHub stars: 3
- Drnoise Benchmarking Deep Research Agents In MisleDerived from arXiv:2607.17291 - DRNOISE: Benchmarking Deep Research Agents in Misleading Evidence EnvironmentsVotes: 0GitHub stars: 3
- Drop Clause Enhancing Performance Interpretability And Robustness Of The Tsetlin Machine**arXiv ID:** 2105.14506 **Authors:** Jivitesh Sharma, Rohan Yadav, Ole-Christoffer Granmo, Lei Jiao **Published:** 2021-05-30T11:29:49Z **Abstract:** In this article, we introduce a novel variant of the Tsetlin machine (TM) that randomly drops clauses, the key learning elements of a TM. In effect, TM with drop clause ignores a random selection of the clauses in each epoch, selected according to a predefined probability. In this way, additional stochasticity is introduced in the learning phas...Votes: 0GitHub stars: 3
- Early Stopping By Correlating Online Indicators In Neural Networks**arXiv ID:** 2402.02513 **Authors:** Manuel Vilares Ferro, Yerai Doval Mosquera, Francisco J. Ribadas Pena, Victor M. Darriba Bilbao **Published:** 2024-02-04T14:57:20Z **Abstract:** In order to minimize the generalization error in neural networks, a novel technique to identify overfitting phenomena when training the learner is formally introduced. This enables support of a reliable and trustworthy early stopping condition, thus improving the predictive power of that type of modeling. Our pr...Votes: 0GitHub stars: 3
- Engineering Hybrid Physicsinformed Neural Networks For Nextgeneration Electricity Systems A Stateoftheart Review**arXiv ID:** 2605.21903 **Authors:** Joseph Nyangon **Published:** 2026-05-21T02:25:05Z **Abstract:** The integration of machine learning with domain-specific physics is transforming the design, monitoring, and control of electricity systems, where data scarcity, limited interpretability, and the need to enforce physical laws constrain purely data-driven models. Physics-informed machine learning (PIML) addresses these limitations by embedding governing equations directly into the learning pr...Votes: 0GitHub stars: 3
- Enhancing Physicsinformed Neural Networks With Domainaware Fourier Features Towards Improved Performance And Interpretable Results**arXiv ID:** 2603.02948 **Authors:** Alberto Miño Calero, Luis Salamanca, Konstantinos E. Tatsis **Published:** 2026-03-03T12:55:53Z **Abstract:** Physics-Informed Neural Networks (PINNs) incorporate physics into neural networks by embedding partial differential equations (PDEs) into their loss function. Despite their success in learning the underlying physics, PINN models remain difficult to train and interpret. In this work, a novel modeling approach is proposed, which relies on the use of...Votes: 0GitHub stars: 3
- Episodic Memory Theory For The Mechanistic Interpretation Of Recurrent Neural Networks**arXiv ID:** 2310.02430 **Authors:** Arjun Karuvally, Peter Delmastro, Hava T. Siegelmann **Published:** 2023-10-03T20:52:37Z **Abstract:** Understanding the intricate operations of Recurrent Neural Networks (RNNs) mechanistically is pivotal for advancing their capabilities and applications. In this pursuit, we propose the Episodic Memory Theory (EMT), illustrating that RNNs can be conceptualized as discrete-time analogs of the recently proposed General Sequential Episodic Memory Model. To s...Votes: 0GitHub stars: 3
- Evaluating Fairness In ChatgptSkill for AI agent capabilitiesVotes: 0GitHub stars: 3
- Fairness Aware System OptimizationFairness-aware strategic design for shared infrastructure systems using bi-objective trajectory-based optimization. Balances revenue maximization with service equity through max-min fairness and service-rate disparity paradigms. Activation: fairness-aware optimization, system design, multi-objective optimization, service equity, shared resource systems, Pareto frontier analysis.Votes: 0GitHub stars: 3
- Fairness In Generative Modeling**arXiv ID:** 2210.03517 **Authors:** Mariia Zameshina, Olivier Teytaud, Fabien Teytaud, Vlad Hosu, Nathanael Carraz, Laurent Najman, Markus Wagner **Published:** 2022-10-06T06:53:02Z **Abstract:** We design general-purpose algorithms for addressing fairness issues and mode collapse in generative modeling. More precisely, to design fair algorithms for as many sensitive variables as possible, including variables we might not be aware of, we assume no prior knowledge of sensitive variables: our...Votes: 0GitHub stars: 3
- Forgetbench Benchmarking Forgetting Dynamics Of LoForgetBench: Benchmarking Forgetting Dynamics of Long-Term Parametric Memory in Language ModelsVotes: 0GitHub stars: 3
- Fuzzy Recurrent Stochastic Configuration Networks For Industrial Data Analytics**arXiv ID:** 2407.11038 **Authors:** Dianhui Wang, Gang Dang **Published:** 2024-07-06T01:40:31Z **Abstract:** This paper presents a novel neuro-fuzzy model, termed fuzzy recurrent stochastic configuration networks (F-RSCNs), for industrial data analytics. Unlike the original recurrent stochastic configuration network (RSCN), the proposed F-RSCN is constructed by multiple sub-reservoirs, and each sub-reservoir is associated with a Takagi-Sugeno-Kang (TSK) fuzzy rule. Through this hybrid fram...Votes: 0GitHub stars: 3
- Glasswing Vulnerability DiscoveryAI-powered vulnerability discovery methodology from Anthropic's Project Glasswing - using frontier models (Mythos Preview) to find critical security vulnerabilities in open-source and enterprise software.Votes: 0GitHub stars: 3
- Gnn Drug Toxicity ExplainabilityGNN-based drug toxicity prediction explainability methodology with Gap Taxonomy (GAP-1 to GAP-4) for systematic analysis of explainability limitations. Uses GNNExplainer on MPNN models trained on Tox21 benchmark.Votes: 0GitHub stars: 3
- Hamqasbench Hamiltonian Qas BenchmarkHamiltonian-informed diagnostic benchmark for Quantum Architecture Search (QAS). Organizes molecules into structural tiers via Pauli operator fingerprints, computational basis representation, and ground-state entanglement. Detects failure modes invisible to energy-only metrics: over-parameterization, eigenstate commitment, representation bottleneck, topology-induced routing failure. Activation: QAS benchmark, Hamiltonian structure, quantum architecture search evaluation, circuit structure ana...Votes: 0GitHub stars: 3
- Iit Fep Maxcaliber BridgeMaximum-Caliber Deviation framework bridging Integrated Information Theory (IIT) with the Free Energy Principle (FEP). Defines information as deviation from constrained maximum-caliber path ensembles, re-derives IIT cause/effect repertoires from variational principles, connects to active inference. Activation: iit fep bridge, maximum caliber, integrated information theory free energy, consciousness framework, constrained entropy maximization.Votes: 0GitHub stars: 3
- Institutional Red Teaming Deployment Rules Not Just Models Causally Shape MultiWe introduce institutional red-teaming, an evaluation methodology for testing deployment rules in multi-agent AI: hold the agents, objectives, and task state fixed, vary only one rule, and attribute t. Based on arXiv:2607.07695.Votes: 0GitHub stars: 3
- Interpretable Neuralsymbolic Concept Reasoning**arXiv ID:** 2304.14068 **Authors:** Pietro Barbiero, Gabriele Ciravegna, Francesco Giannini, Mateo Espinosa Zarlenga, Lucie Charlotte Magister, Alberto Tonda, Pietro Lio', Frederic Precioso, Mateja Jamnik, Giuseppe Marra **Published:** 2023-04-27T09:58:15Z **Abstract:** Deep learning methods are highly accurate, yet their opaque decision process prevents them from earning full human trust. Concept-based models aim to address this issue by learning tasks based on a set of human-understandabl...Votes: 0GitHub stars: 3