Skip to content
Back to skills

Ai Datacenter Networking

ASecurity

Lay out the network for distributed training so collectives run on the fastest link that spans them, using NVLink, InfiniBand, and topology-aware placement. Use when a training job spans multiple GPUs or nodes and interconnect, not compute, is capping throughput.

  • 7 stars
  • 0 votes
  • 0 copies
  • 1 view
  • Added September 5, 2026
ai-agentsrustgonode

Security analysis

A100/100

Scanned September 5, 2026

npx -y skills add Amey-Thakur/AI-SKILLS --skill ai-datacenter-networking --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Ai Datacenter Networking?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Ai Datacenter Networking
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/amey-thakur-ai-datacenter-networking/badge)](https://www.skillsdirectory.com/skills/amey-thakur-ai-datacenter-networking)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: ai-datacenter-networking
description: Lay out the network for distributed training so collectives run on the fastest link that spans them, using NVLink, InfiniBand, and topology-aware placement. Use when a training job spans multiple GPUs or nodes and interconnect, not compute, is capping throughput.
---

# AI datacenter networking

Training at scale is a communication problem wearing a compute costume. Every
gradient all-reduce and every activation exchange rides a link, and the fabric
has tiers that differ by more than 10x in bandwidth. Place a collective on the
wrong tier and the GPUs sit idle waiting on the wire, so the job is to match
each communication pattern to the fastest link that contains its group.

## Method

1. **Inventory the fabric tiers by bandwidth.** Inside a node, NVLink and
   NVSwitch give roughly 900 GB/s per H100 all-to-all. Between nodes, InfiniBand
   NDR runs 400 Gb/s per port, about 50 GB/s. That gap, near 18x, is why intra
   and inter-node traffic must be treated as different budgets, not one network.
2. **Bind each collective to the tier that spans its group.** Keep the
   tensor-parallel all-reduce, the chattiest collective, inside one NVLink node
   of 8 GPUs. Let data-parallel gradient all-reduce and pipeline point-to-point
   cross nodes on InfiniBand, where their lower frequency tolerates the slower
   link.
3. **Demand a rail-optimized topology.** In a rail design each GPU reaches the
   leaf switch through its own NIC on a fixed rail, so same-rank GPUs across
   nodes talk without crossing the spine. Confirm one NIC per GPU and rail
   alignment; a shared or misrailed NIC halves effective bandwidth under load.
4. **Turn on GPUDirect RDMA and point NCCL at the right HCAs.** Set
   `NCCL_IB_HCA` to the InfiniBand adapters, `NCCL_NET_GDR_LEVEL` to allow
   GPU-to-NIC DMA, and supply a `NCCL_TOPO_FILE` so NCCL builds rings and trees
   that follow the real wiring instead of guessing across a PCIe hop.
5. **Place ranks with a topology-aware scheduler.** Use SLURM block placement or
   the cluster's topology plugin to pack a job into adjacent nodes under one
   leaf switch. Scattering 64 ranks across a congested spine adds hops and
   queueing that no NCCL tuning recovers.
6. **Prefer in-network reduction and adaptive routing where the fabric offers
   it.** SHARP offloads all-reduce into the switch ASIC, cutting data that
   crosses the wire. Adaptive routing spreads flows so a single hot link does
   not stall a collective. Enable both when the hardware supports them.
7. **Measure busbw before you trust the layout.** Run `all_reduce_perf` from
   nccl-tests and read bus bandwidth, not raw algbw. If NVLink groups miss ~80
   percent of peak or IB groups fall well under line rate, a rank crossed the
   wrong tier and the plan is wrong on paper.

## Signals

- Do tensor-parallel groups stay within a single NVLink node in the placement?
- Does nccl-tests busbw land near peak for both the NVLink and InfiniBand
  groups?
- Is every GPU mapped to its own rail-aligned NIC with GPUDirect RDMA active?
- Did the scheduler pack ranks under shared leaf switches instead of the spine?

## Boundaries

This covers wiring collectives to the fabric, not choosing how to split the
model across devices, which is model-parallelism. Sizing GPUs and interconnect
before purchase is gpu-cost-planning. The optimal NCCL algorithm and tree shape
are cluster-specific and come from measurement, not a fixed recipe.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…