Skip to content
Back to skills

Infinitevl Linear Sparse Attention

ASecurity

Merge sliding window and linear attention (Gated DeltaNet) for unlimited VLM inputs with 3.6× speedup. InfiniteVL handles video understanding at 24 FPS with constant memory—ideal when context length must scale without quadratic overhead.

  • 6 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 9, 2026
code-qualitypythonperformance

Security analysis

A100/100

Scanned September 9, 2026

npx -y skills add ADu2021/skillXiv --skill infinitevl-linear-sparse-attention --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Infinitevl Linear Sparse Attention?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Infinitevl Linear Sparse Attention
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/adu2021-infinitevl-linear-sparse-attention/badge)](https://www.skillsdirectory.com/skills/adu2021-infinitevl-linear-sparse-attention)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: infinitevl-linear-sparse-attention
title: "InfiniteVL: Synergizing Linear and Sparse Attention for Highly-Efficient, Unlimited-Input Vision-Language Models"
version: 0.0.2
engine: skillxiv-v0.0.2-claude-opus-4.6
license: MIT
url: https://arxiv.org/abs/2512.08829
keywords: [vision-language models, linear attention, sparse attention, long sequence, efficient inference]
description: "Merge sliding window and linear attention (Gated DeltaNet) for unlimited VLM inputs with 3.6× speedup. InfiniteVL handles video understanding at 24 FPS with constant memory—ideal when context length must scale without quadratic overhead."
---

## Overview

InfiniteVL combines sparse window-based attention for local detail with linear attention mechanisms for global efficiency, enabling unlimited input handling without quadratic complexity growth or expanding KV cache issues.

## When to Use

- Vision-language models processing variable-length inputs
- Long video understanding requiring stable 24 FPS performance
- OCR and information-intensive vision tasks
- Unlimited input sequences without memory explosion
- Need for 3.6× speedup over transformer baselines

## When NOT to Use

- Short-context tasks benefiting from standard transformers
- Scenarios where window size is sufficient
- Information-intensive tasks preferring dense attention

## Core Technique

Hybrid attention architecture synergizing sparse and linear mechanisms:

```python
# InfiniteVL: Hybrid linear + sparse attention
class InfiniteVLAttention(nn.Module):
    def __init__(self, dim, window_size=1024):
        super().__init__()
        self.window_attn = SlidingWindowAttention(window_size)
        self.linear_attn = GatedDeltaNetAttention(dim)

    def forward(self, query, key, value, seq_len):
        """Synergize sparse and linear attention."""
        # Sparse attention: sliding window for local context
        local_output = self.window_attn(query, key, value)

        # Linear attention: global efficiency via Gated DeltaNet
        global_output = self.linear_attn(query, key, value)

        # Combine: local detail + global structure
        output = local_output + 0.3 * global_output

        return output
```

## Key Results

- 3.6× inference speedup
- Constant latency and memory with sequence length
- 24 FPS video streaming performance
- Matches leading transformer-based VLMs

## References

- Original paper: https://arxiv.org/abs/2512.08829
- Focus: Efficient long-sequence VLMs
- Domain: Vision-language models, attention mechanisms

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…