Skip to content
Back to skills

Vector

ASecurity

Vector is a high-performance observability data pipeline written in Rust and maintained by Datadog. It collects, transforms and routes logs, metrics and traces from any source to any destination using TOML or YAML configs and the VRL remap language. Use it to replace Logstash, Fluentd or Filebeat, filter and sample noisy logs, archive to S3, or write and debug vector.toml, VRL and `vector test` unit tests.

  • 142 stars
  • 0 votes
  • 0 copies
  • 3 views
  • Added September 6, 2026
devopsrustgoshellbashnodedockerkubernetesawsgitapi

Works with

  • terminal
  • api

Security analysis

A100/100

Pro scans all 2 files and shows the line behind each finding

Scanned October 4, 2026

npx -y skills add TerminalSkills/skills --skill vector --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Vector?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Vector
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/terminalskills-vector/badge)](https://www.skillsdirectory.com/skills/terminalskills-vector)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: vector
description: >-
  Vector is a high-performance observability data pipeline written in Rust and
  maintained by Datadog. It collects, transforms and routes logs, metrics and
  traces from any source to any destination using TOML or YAML configs and the
  VRL remap language. Use it to replace Logstash, Fluentd or Filebeat, filter
  and sample noisy logs, archive to S3, or write and debug vector.toml, VRL and
  `vector test` unit tests.
license: Apache-2.0
compatibility: Vector 0.58 or newer (config shown was validated with 0.58.0); Linux, macOS, Windows, Docker or Kubernetes
metadata:
  author: terminal-skills
  version: 1.1.0
  category: devops
  repository: https://github.com/vectordotdev/vector
  tags:
  - log-pipeline
  - data-pipeline
  - observability
  - rust
  - transform
---

# Vector — High-Performance Observability Data Pipeline

## Overview

Vector reads data from **sources**, reshapes it with **transforms** and delivers it to **sinks**. Each component has a name and an `inputs` list that wires it to upstream components. It runs as an agent on every host (DaemonSet in Kubernetes) or as a central aggregator, and ships as a single static binary. Configuration can be TOML, YAML or JSON; this skill uses TOML.

Changes in recent releases that older tutorials get wrong:

- **0.57**: environment-variable interpolation (`${ES_PASSWORD}`) in config files is **off by default**. Re-enable it with `--dangerously-allow-env-var-interpolation` or `VECTOR_DANGEROUSLY_ALLOW_ENV_VAR_INTERPOLATION=true`, run the file through `envsubst` first, or use a secrets backend.
- **0.57**: sinks that accept `{{ field }}` templates (S3, Elasticsearch, Kafka, HTTP, Loki, file and others) refuse templates with no literal prefix, such as `topic = "{{ x }}"`. Write `topic = "logs-{{ x }}"`, or set `dangerously_allow_unconfined_template_resolution = true` only if you accept the injection risk. `vector validate` reports these at startup.
- Elasticsearch `auth` needs `auth.strategy = "basic"`; without it validation fails with "missing field `strategy`".
- The `http` sink has no `condition` option; route errors to it through a `filter` transform, and keep errors out of sampling with `exclude` on `sample`.
- Latest release when checked: 0.58.0 (26 August 2026).

## Instructions

### Install

Prefer a package manager or container image:

```bash
brew install vector                       # macOS
helm repo add vector https://helm.vector.dev
helm install vector vector/vector         # Kubernetes
docker run --rm -v "$PWD/vector.toml:/etc/vector/vector.toml:ro" \
  timberio/vector:0.58.0-alpine --config /etc/vector/vector.toml
```

For apt, yum and archive installs follow https://vector.dev/docs/setup/installation/. If you download an archive from packages.timber.io, verify its published checksum before unpacking, and avoid piping the install script into a shell.

### Configuration

```toml
# vector.toml — collect, filter, sample and route logs and metrics
data_dir = "/var/lib/vector"

# --- Sources ---
[sources.app_logs]
type = "file"
include = ["/var/log/app/*.log"]
read_from = "end"                     # "beginning" re-reads old files on first start

[sources.http_logs]                   # apps POST JSON here
type = "http_server"
address = "127.0.0.1:8686"
decoding.codec = "json"

[sources.host_metrics]
type = "host_metrics"
collectors = ["cpu", "memory", "disk", "network"]
scrape_interval_secs = 15

[sources.otel]                        # OpenTelemetry: outputs are otel.logs / otel.metrics / otel.traces
type = "opentelemetry"
grpc.address = "127.0.0.1:4317"
http.address = "127.0.0.1:4318"

# --- Transforms ---
[transforms.parse_json]
type = "remap"
inputs = ["app_logs"]
source = '''
  . = parse_json!(.message)
  .environment = get_env_var("ENVIRONMENT") ?? "production"
  .timestamp = parse_timestamp!(.timestamp, format: "%Y-%m-%dT%H:%M:%S%.fZ")
  if exists(.email) {
    .email = redact(string!(.email), filters: [r'\S+@\S+'])
  }
'''

[transforms.filter_noise]
type = "filter"
inputs = ["parse_json"]
condition = '''
  !includes(["GET /health", "GET /ready", "GET /metrics"], .request) && .level != "debug"
'''

[transforms.sample_info]              # keeps 1 in 10 events, never drops errors or warnings
type = "sample"
inputs = ["filter_noise"]
rate = 10
exclude = '.level == "error" || .level == "warn"'

[transforms.errors_only]
type = "filter"
inputs = ["filter_noise"]
condition = '.level == "error"'

[transforms.aggregate_metrics]
type = "aggregate"
inputs = ["host_metrics"]
interval_ms = 60000

# --- Sinks ---
[sinks.elasticsearch]
type = "elasticsearch"
inputs = ["sample_info"]
endpoints = ["https://es.internal.shopwave.io:9200"]
bulk.index = "logs-%Y-%m-%d"
auth.strategy = "basic"
auth.user = "${ES_USER}"              # needs env-var interpolation enabled (see Overview)
auth.password = "${ES_PASSWORD}"
compression = "gzip"
batch.max_bytes = 10485760
batch.timeout_secs = 5
buffer.type = "disk"                  # survives destination outages and restarts
buffer.max_size = 1073741824
buffer.when_full = "block"

[sinks.s3_archive]
type = "aws_s3"
inputs = ["sample_info"]
bucket = "shopwave-logs-archive"
region = "us-east-1"
key_prefix = "logs/{{ service }}/year=%Y/month=%m/day=%d/"   # literal prefix keeps 0.57+ confinement happy
compression = "gzip"
encoding.codec = "json"
batch.max_bytes = 104857600
batch.timeout_secs = 300

[sinks.prometheus]
type = "prometheus_exporter"
inputs = ["aggregate_metrics"]
address = "0.0.0.0:9598"

[sinks.slack_errors]
type = "http"
inputs = ["errors_only"]
uri = "${SLACK_WEBHOOK_URL}"
method = "post"
encoding.codec = "json"
batch.max_events = 1
request.rate_limit_duration_secs = 1
request.rate_limit_num = 5
```

Validate, then run:

```bash
vector validate --no-environment vector.toml      # add --dangerously-allow-env-var-interpolation if the file uses ${VAR}
vector --config vector.toml
```

`--no-environment` skips network health checks, so it works offline; drop it in CI to also check that destinations are reachable. Unused sources only produce warnings.

### VRL (Vector Remap Language)

VRL is compiled and type-checked: fallible calls must be handled with `!` (abort the event on error) or `?? default`.

```coffee
. = parse_json!(.message)

if starts_with(string!(.message), "AUDIT:") {
  .route = "audit"
} else if (to_int(.status_code) ?? 0) >= 500 {
  .route = "error"
} else {
  .route = "general"
}

.duration_ms = to_float(.duration_ms) ?? 0.0
.user_id = del(.user.id)
del(.user)
```

Try snippets interactively with `vector vrl`, or watch live events with `vector tap` and `vector top` against a running instance.

### Unit tests

```toml
[[tests]]
name = "normalizes warning and redacts email"
[[tests.inputs]]
insert_at = "parse_json"
type = "log"
[tests.inputs.log_fields]
message = '{"level":"warning","status_code":"503","email":"dana@shopwave.io"}'
[[tests.outputs]]
extract_from = "parse_json"
[[tests.outputs.conditions]]
type = "vrl"
source = '''
  assert_eq!(.severity, "warn")
  assert!(!contains(string!(.email), "dana"))
'''
```

Put the `[[tests]]` next to the transforms they cover and run `vector test vector.toml`. Output: `test normalizes warning and redacts email ... passed`.

## Examples

### Example 1: Cut Elasticsearch cost for a Node.js API

**User request:** "Our Node API logs JSON to /var/log/api/*.log and Elasticsearch costs too much. Drop health checks and debug, keep all errors, sample the rest."

Create `vector.toml` from the configuration above, with `include = ["/var/log/api/*.log"]`. Run `vector validate --no-environment vector.toml`; it prints `Validated` for transforms and sinks. Start with `vector --config vector.toml` and confirm with `vector top` that `filter_noise` receives more events than `sample_info` emits. Errors and warnings all reach Elasticsearch; about one in ten info events does.

### Example 2: Fix a failing VRL remap

**User request:** "Vector says `unhandled fallible assignment` on `.timestamp = parse_timestamp(.ts, \"%+\")`."

`parse_timestamp` can fail, so VRL refuses to compile it. Change the line to `.timestamp = parse_timestamp!(.ts, format: "%+")` to drop events with a bad timestamp, or `parse_timestamp(.ts, format: "%+") ?? now()` to keep them. Add a `[[tests]]` case with a malformed `ts` and run `vector test vector.toml` to lock in the behaviour.

## Guidelines

- Pin the image or package version; upgrade one minor at a time and read the upgrade guide, since 0.57 changed env-var and template handling.
- Never put secrets in the config file. Use a secrets backend, or env vars with interpolation explicitly enabled, and keep the file out of Git.
- Filter noise and sample high-volume info logs before sinks; keep 100% of errors and warnings.
- Archive everything cheaply to S3 and send only recent or important data to Elasticsearch.
- Enable disk buffers on sinks that must not lose data during outages; the default buffer is in memory.
- Bind source addresses to `127.0.0.1` unless remote hosts must reach them; `0.0.0.0` on syslog or HTTP exposes the port.
- Prefer VRL over regex-heavy chains and test every remap with `vector test` before deploying.

Files in this skill

  • SKILL.md7.4 KB
  • _scores.json1.8 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…