Skip to content
Back to skills

Bigquery Optimization

ASecurity

Provides workflows to optimize BigQuery environments (capacity planning, editions), storage assets (partitioning, clustering, storage lifecycles, billing models), and SQL queries. Use when optimizing cost, modeling Edition migrations, rightsizing reservations, evaluating logical vs. physical storage, designing table partitioning/clustering, generating table DDL, migrating unpartitioned tables, managing partition expiration, or optimizing individual SQL queries. Do not use for raw usage repor...

  • 4 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added October 1, 2026
ai-agentsgoshellbashsqlrailsapi

Works with

  • cli
  • api

Security analysis

A100/100

Pro scans all 6 files and shows the line behind each finding

Scanned October 1, 2026

npx -y skills add nuroctane/nur-cli --skill bigquery-optimization --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Bigquery Optimization?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Bigquery Optimization
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/nuroctane-bigquery-optimization/badge)](https://www.skillsdirectory.com/skills/nuroctane-bigquery-optimization)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: bigquery-optimization
metadata:
  version: "1.0.0"
  category: BigDataAndAnalytics
description: >-
  Provides workflows to optimize BigQuery environments (capacity planning,
  editions), storage assets (partitioning, clustering, storage lifecycles,
  billing models), and SQL queries. Use when optimizing cost, modeling Edition
  migrations, rightsizing reservations, evaluating logical vs. physical storage,
  designing table partitioning/clustering, generating table DDL, migrating
  unpartitioned tables, managing partition expiration, or optimizing individual
  SQL queries.

  Do not use for raw usage reporting (use bigquery-observability), query
  execution plan analysis, error troubleshooting, or diagnosing why a specific
  job was slow (use bigquery-troubleshooting).
---

# BigQuery Optimization Workflow

## Prerequisites & Environment Setup

Before executing optimization analyses, evaluating editions, or applying DDL
modifications:

1.  **Google Cloud SDK**: Ensure the
    [Google Cloud SDK](https://cloud.google.com/sdk/docs/install) is installed
    and configured.
2.  **Project Selection**: Set the active Google Cloud project:

    ```bash
    gcloud config set project {project_id}
    ```

3.  **API Enablement**: Ensure BigQuery and BigQuery Reservation APIs are
    enabled:

    ```bash
    gcloud services enable \
        bigquery.googleapis.com bigqueryreservation.googleapis.com
    ```

4.  **Authentication**: Authenticate the environment:
    *   CLI tools and `bq` commands: `gcloud auth login`
    *   SDKs and automation: `gcloud auth application-default login`
    *   Service accounts: Set
        `GOOGLE_APPLICATION_CREDENTIALS="/path/to/key.json"`

5.  **Billing & IAM Roles**:
    *   Verify an active Google Cloud Billing account is attached to
        `{project_id}`.
    *   Ensure appropriate IAM roles:
        *   `roles/bigquery.admin` or `roles/bigquery.resourceAdmin`:
            Reservation and capacity commitment management.
        *   `roles/bigquery.dataEditor` or `roles/bigquery.admin`: Modifying
            table schemas, partitioning, clustering, and storage billing
            models.
        *   `roles/bigquery.jobUser`: Running evaluation queries.

6.  **Companion Skills Installation**:
    This skill is part of a 3-pillar operations suite (`bigquery-observability`,
    `bigquery-optimization`, `bigquery-troubleshooting`). If any companion skill
    is not yet installed in your environment, install the full suite:

    ```bash
    npx skills add google/skills --skill bigquery-observability --skill bigquery-optimization --skill bigquery-troubleshooting
    ```

    *(If `bigquery-observability` is not installed, use the self-contained
    baseline formulas and query templates provided directly in the reference
    sections below).*

## Workflows

Determine the optimization focus of the user's request and follow the relevant
workflow:

-   **Telemetry & Observability Baseline:** For direct raw usage telemetry,
    `INFORMATION_SCHEMA` queries, and baseline metric calculations, consult
    [bigquery-observability](../bigquery-observability)
    (`bigquery_observability`). If the `bigquery-observability` companion skill
    is not available in the active environment, all optimization guidelines, DDL
    templates, and decision models across this skill and its reference guides
    are fully self-contained.
-   **Capacity & Editions Modeling:** Evaluate the cost-efficiency of migrating
    workloads from On-Demand to Editions, as well as rightsizing active Edition
    reservations, baseline commitments, and autoscaling caps.
    *   *Instructions:* Read `references/capacity_planning_editions.md` to
        provide deep links to BigQuery's built-in recommendation UIs (e.g., Slot
        Estimator) and guide the user through UI navigation: 1. navigate to the
        Slot Estimator tab, 2. select 'On-Demand' as the source to analyze
        historical query volume, and 3. review the Cost-Optimized
        Recommendations and Slot Usage Chart.
-   **Table & Storage Optimization:** Optimize storage costs from a billing
    model, physical layout, and lifecycle perspective.
    *   *Billing Architecture:* Read `references/storage_billing_models.md` for
        guidance on evaluating aggregate compression ratios (e.g. >2:1 threshold
        in US) to recommend Physical vs. Logical billing, noting that the
        break-even ratio depends on specific regional rates and custom
        enterprise contracts. When providing `TABLE_STORAGE` queries, always
        scope with `WHERE table_schema = '{dataset_id}'`, use the regional
        dataset view, and warn that 0 rows indicates a region mismatch or lack
        of native tables rather than zero billable usage.
    *   *Partitioning & Clustering Strategy:* Read
        `references/table_partitioning_clustering.md` to generate production DDL
        templates (CREATE TABLE, CTAS migrations for unpartitioned tables, and
        modifying clustering specifications), enforce pruning with
        `require_partition_filter = true`, and manage partition limits (up to
        10,000 partitions/table).
    *   *Lifecycle Management:* Read
        `references/storage_lifecycle_management.md` to pinpoint inactive data
        and define precise Time-to-Live (TTL) partition expirations, dataset
        expirations, and Time Travel window reductions.
-   **SQL Optimization:** Optimize individual SQL queries to reduce slot-time
    and the amount of data read.
    *   *Instructions:* Follow the instructions in
        `references/sql_optimization.md` to provide recommendations to the user
        on how to rewrite their SQL query to reduce slot-time and the amount of
        data read.

## Execution Guardrails

-   **Terminology & Cost Framing:** Never promise or guarantee "cost-reduction"
    or "reducing expenditure." Always frame recommendations using the
    terminology **"optimizing your bill"** or **"improving cost-efficiency."**
-   **Explicit Scope Framing & Region Resolution:** Always state the target
    `project_id` and `region` at the very top of your response so the user
    immediately knows the exact scope being evaluated. Follow this 3-tier
    resolution hierarchy:
    1.  *Explicit Region:* Use the region specified in the user's prompt (e.g.,
        `europe-west1`).
    2.  *Contextual Region:* Resolve the region from the specific dataset or
        resource mentioned in the context.
    3.  *Unspecified Fallback:* Default to `us` / `region-us`, explicitly state
        that `us` was assumed as the default, and instruct the user to
        substitute their region if their resources reside elsewhere. *Region
        Formatting:* In Cloud Console deep links, use the region identifier
        directly (e.g., `region=us`, `region=europe-west1`). In SQL queries
        against `INFORMATION_SCHEMA`, use the regional dataset qualifier (e.g.,
        `region-us`, `region-europe-west1`).
-   **Zero-Row Result Guard:** If querying `TABLE_STORAGE` with `WHERE
    table_schema = '{dataset_id}'` returns 0 rows, do not proceed with an empty
    or zero-usage evaluation. Treat this as an indicator that the dataset may
    reside in a different region or have no native tables; stop and prompt the
    user to confirm the dataset's regional location.
-   **Populate Concrete Parameters:** When generating URLs and SQL queries,
    always substitute known `project_id` and `region` values directly into the
    code and links. Never leave literal `{project_id}` or `{location}`
    placeholders for the user to manually edit.
-   **No Autonomous Purchasing or Financial Mutations:** Never provide the user
    with executable scripts (e.g., `gcloud` or `bq` shell commands like `bq
    update --storage_billing_model=...`) designed to autonomously purchase
    annual commitments, alter edition tier bindings, or mutate storage billing
    models. Always guide the user to execute commitment purchases, reservation
    changes, and storage billing model updates manually via the Cloud Console
    UI.

Files in this skill

  • SKILL.md7.9 KB
  • references/capacity_planning_editions.md7 KB
  • references/sql_optimization.md17.4 KB
  • references/storage_billing_models.md6.2 KB
  • references/storage_lifecycle_management.md2.2 KB
  • references/table_partitioning_clustering.md12.2 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…