---
name: honeycomb
description: Use when instrumenting applications and infrastructure to send observability data (traces, logs, metrics), querying telemetry to investigate issues, setting up alerts and SLOs, managing data volume, and building dashboards for monitoring systems.
metadata:
    mintlify-proj: honeycomb
    version: "1.0"
---

# Honeycomb Skill

## Product Summary

Honeycomb is an observability platform for understanding complex distributed systems through high-cardinality event data. Unlike traditional tools that pre-aggregate metrics or limit queryable fields, Honeycomb stores raw telemetry as structured events and indexes every field automatically, enabling arbitrary queries at interactive speed. Agents use Honeycomb to instrument applications (via OpenTelemetry SDKs), send infrastructure data (via collectors), query traces/logs/metrics to investigate issues, set up alerts (Triggers and SLOs), and build dashboards (Boards).

**Key files and concepts:**
- **API Keys**: Ingest Keys (send data), Configuration Keys (manage resources), Management Keys (team-level operations). Pass via `X-Honeycomb-Team` header or `Authorization: Bearer` for Management Keys.
- **Resource hierarchy**: Teams → Environments → Datasets. Datasets are created automatically when you send data with a service name.
- **Telemetry signals**: Traces (spans with parent-child relationships), Logs (structured events), Metrics (time-series data in dedicated datasets).
- **Query Builder**: Filter (WHERE), group (GROUP BY), aggregate (SELECT), sort (ORDER BY), limit results (LIMIT), filter aggregates (HAVING).
- **Primary docs**: https://docs.honeycomb.io

## When to Use

Reach for this skill when:
- **Instrumenting code**: Agent needs to add OpenTelemetry SDKs to applications or configure collectors for infrastructure.
- **Querying data**: Agent must investigate issues by building queries, exploring traces, analyzing logs, or correlating metrics.
- **Setting up alerts**: Agent creates Triggers (threshold-based) or SLOs (reliability targets with error budgets) and configures notifications.
- **Managing data**: Agent needs to sample data, filter events, adjust metrics granularity, or organize datasets/environments.
- **Building dashboards**: Agent creates or updates Boards with queries, SLOs, and visualizations.
- **Configuring resources**: Agent manages API keys, calculated fields, markers, or dataset definitions.
- **Using the API**: Agent makes programmatic requests to create queries, boards, triggers, SLOs, or manage datasets.

## Quick Reference

### API Key Types and Headers

| Key Type | Scope | Use | Header |
|----------|-------|-----|--------|
| Ingest | Environment | Send telemetry | `X-Honeycomb-Team: <key_id><secret>` |
| Configuration | Environment | Manage resources (queries, boards, triggers, SLOs) | `X-Honeycomb-Team: <token>` |
| Management | Team | Manage keys, environments | `Authorization: Bearer <key_id>:<secret>` |

### Query Builder Clauses

| Clause | Purpose | Example |
|--------|---------|---------|
| SELECT | Aggregate calculation | `COUNT`, `P95(duration_ms)`, `AVG(response_time)` |
| WHERE | Filter events | `status_code = 500`, `error exists` |
| GROUP BY | Group results | `service.name`, `http.method` |
| ORDER BY | Sort results | `P95(duration_ms) DESC` |
| LIMIT | Limit rows | `100` (default) |
| HAVING | Filter aggregates | `COUNT > 10` |

### Relational Field Prefixes (WHERE and GROUP BY only)

- `root.` — Root span in trace
- `parent.` — Direct parent span
- `child.` — Single child span
- `any.`, `any2.`, `any3.` — Spans anywhere in trace
- `none.` — Exclude traces containing a span

### Common SELECT Aggregates

`COUNT`, `COUNT_DISTINCT(field)`, `SUM(field)`, `AVG(field)`, `MIN(field)`, `MAX(field)`, `P50(field)`, `P95(field)`, `P99(field)`, `HEATMAP(field)`, `CONCURRENCY`, `RATE_AVG(field)`, `RATE_SUM(field)`

### Data Organization

| Level | Purpose | Notes |
|-------|---------|-------|
| Team | Top-level account | Contains multiple environments |
| Environment | Logical grouping | Group by deployment stage (prod, staging, dev) or regulatory boundary |
| Dataset | Collection of events | Created automatically by service name; max 100 per environment (300 for Enterprise) |

### Calculated Fields (Derived Columns)

Create formulas to transform data at query time. Scope: dataset-level or environment-level. Use in queries like any other field.

Example: `if(duration_ms > 1000, "slow", "fast")`

## Decision Guidance

### When to Use Traces vs. Logs vs. Metrics

| Signal | Use When | Example |
|--------|----------|---------|
| **Traces** | Following a request across services, understanding execution flow, correlating errors with code paths | "Why is this API call slow?" |
| **Logs** | Recording discrete events, errors, state changes, application output | "What errors occurred in the last hour?" |
| **Metrics** | Monitoring infrastructure health, tracking rates/trends, alerting on thresholds | "Is CPU utilization high?" |

### When to Use Triggers vs. SLOs

| Feature | Use When | Example |
|---------|----------|---------|
| **Triggers** | Threshold-based alerts on any query; flexible conditions; immediate notification | Alert when error rate > 5% in last 5 minutes |
| **SLOs** | Reliability targets with error budgets; burn rate tracking; multi-service support | 99.9% of requests succeed within 250ms over 30 days |

### When to Sample Data

| Scenario | Approach |
|----------|----------|
| High-volume, low-cardinality data (e.g., successful requests) | Use head sampling (deterministic or dynamic) in instrumentation |
| Need to keep errors but sample successes | Use Honeycomb Refinery with dynamic sampling rules |
| Already ingesting; need to reduce volume | Use filter processors in OpenTelemetry Collector |

### When to Create Separate Environments

| Scenario | Action |
|----------|--------|
| Production vs. staging data | Separate environments |
| Different regulatory requirements | Separate environments |
| Developer local testing | Shared `dev` environment or per-developer `dev-<name>` |
| Infrastructure metrics supporting multiple stages | Put in one well-known environment (e.g., `prod` or `infra`) |

## Workflow

### 1. Instrument an Application

1. **Choose instrumentation method**: OpenTelemetry SDK (recommended) or libhoney library.
2. **Install SDK**: Follow language-specific docs (Node.js, Python, Go, Java, Ruby, .NET, etc.).
3. **Configure exporter**: Point to Honeycomb endpoint with Ingest Key in `X-Honeycomb-Team` header.
4. **Add custom instrumentation**: Use `tracer.start_as_current_span()` to create spans with attributes.
5. **Set service name**: Ensure `service.name` attribute is set; this becomes the dataset name.
6. **Test**: Send a request and verify data appears in Honeycomb UI within seconds.

### 2. Query Data to Investigate an Issue

1. **Navigate to Query Builder**: Select Query from left menu.
2. **Choose dataset**: Use dataset switcher (top-left) to select service or environment-wide.
3. **Build query**:
   - Add **WHERE** clause to filter (e.g., `error exists`)
   - Add **SELECT** to aggregate (e.g., `COUNT`, `P95(duration_ms)`)
   - Add **GROUP BY** to break down (e.g., `service.name`)
4. **Run query**: Press Shift+Enter or click Run Query.
5. **Explore results**: Click on visualization to view traces, use BubbleUp to identify outliers.
6. **Refine**: Adjust filters, groupings, or time range iteratively.

### 3. Create a Trigger (Threshold Alert)

1. **Build a query** in Query Builder with a **SELECT** clause (e.g., `COUNT`, `P95(duration_ms)`).
2. **Click Trigger icon** on the visualization.
3. **Define condition**: Set threshold (e.g., "greater than 100") and duration/frequency (e.g., "5 minutes, check every 2 minutes").
4. **Add recipients**: Select notification channels (Slack, PagerDuty, email, webhook).
5. **Save**: Name the trigger and optionally add tags.
6. **Verify**: Trigger will fire when condition is met.

### 4. Create an SLO (Reliability Target)

1. **Define SLI**: Create a calculated field or use existing field to measure success (e.g., `duration_ms < 250`).
2. **Navigate to SLOs**: Select SLOs from left menu.
3. **Create SLO**:
   - Select dataset(s) to include.
   - Set SLI expression (e.g., `duration_ms < 250`).
   - Set target (e.g., 99.9%) and window (e.g., 30 days).
4. **Configure burn alerts**: Set burn rate threshold (e.g., alert if burn rate > 2.0).
5. **Add recipients**: Select notification channels.
6. **Save**: Name the SLO and optionally add tags.
7. **Monitor**: View error budget, burn rate, and historical health on SLO detail page.

### 5. Create a Board (Dashboard)

1. **Navigate to Boards**: Select Boards from left menu.
2. **Create Board**: Click "Create Board" or start from a template.
3. **Add panels**:
   - **Query**: Build a query in Query Builder, then add to board.
   - **SLO**: Select an existing SLO.
   - **Text**: Add markdown notes or links.
4. **Customize**: Adjust display (graph only, graph + table, table only).
5. **Save**: Name the board and optionally set as default for a dataset.
6. **Share**: Copy board URL or share with team.

### 6. Manage Data Volume

1. **Identify high-volume sources**: Query `COUNT` grouped by `service.name` or `http.status_code`.
2. **Choose strategy**:
   - **Head sampling**: Add `sample_rate` field in instrumentation (e.g., 1 in 100 for status 200).
   - **Tail sampling**: Deploy Honeycomb Refinery with dynamic sampling rules.
   - **Filtering**: Use OpenTelemetry Collector filter processor to drop unwanted events.
3. **Implement**: Update instrumentation, redeploy, and monitor data volume.
4. **Verify**: Check usage dashboard to confirm reduction.

## Common Gotchas

- **API key scope mismatch**: Using an Ingest Key to manage resources or a Configuration Key to send data will fail. Match key type to operation.
- **Missing service.name**: If instrumentation doesn't set `service.name`, data goes to a default dataset. Always explicitly set service name.
- **Trace spans in different environments**: Traces only render if all spans are in the same environment. Ensure all services send to the same environment.
- **Calculated field scope confusion**: Dataset-level calculated fields only work within that dataset. Use environment-level fields for cross-dataset queries.
- **SLO with low event volume**: SLOs are most effective with high volume. Low-volume SLOs may have noisy burn rates.
- **Sampling without sample_rate field**: If you sample events but don't include `sample_rate`, query results will undercount. Always include the field.
- **Runaway schema from dynamic field names**: Never generate field names dynamically (e.g., from timestamps or user input). This creates unbounded cardinality and throttles new columns.
- **Trigger frequency too high**: Setting trigger frequency to 1 minute on high-volume data can cause alert fatigue. Start with 5-minute frequency.
- **Forgetting to set error field in dataset definition**: Error highlighting in traces relies on the error field being configured. Set it in dataset settings.
- **Querying across incompatible datasets**: Logs and metrics datasets have different query syntax. Stick to one signal type per query.

## Verification Checklist

Before submitting work:

- [ ] **Data is flowing**: Verify events appear in Honeycomb UI within seconds of instrumentation.
- [ ] **Service name is set**: Check that `service.name` attribute is present on all spans/events.
- [ ] **API key has correct scope**: Confirm Ingest Key for sending data, Configuration Key for managing resources.
- [ ] **Query returns results**: Run query and confirm it matches expected data (not empty or too broad).
- [ ] **Trigger/SLO fires correctly**: Test by injecting a failure or manually triggering condition; verify notification is sent.
- [ ] **Calculated field syntax is valid**: Check expression for typos and correct function names.
- [ ] **Sampling is configured**: If using sampling, verify `sample_rate` field is present and Honeycomb scales counts correctly.
- [ ] **Board displays correctly**: Verify all panels load, visualizations render, and time range is appropriate.
- [ ] **Permissions are set**: Confirm team members can access datasets, boards, and alerts they need.
- [ ] **Documentation is clear**: Add descriptions to datasets, SLOs, and boards so team understands purpose.

## Resources

**Comprehensive navigation**: https://docs.honeycomb.io/llms.txt

**Critical documentation pages**:
1. [Introduction to Honeycomb](https://docs.honeycomb.io/get-started/honeycomb/introduction) — Core concepts, event model, signals.
2. [API Authentication](https://docs.honeycomb.io/api/authentication) — API key types, scopes, headers.
3. [Build a Query](https://docs.honeycomb.io/investigate/query/build) — Query Builder reference, clauses, relational fields, aggregates.

---

> For additional documentation and navigation, see: https://docs.honeycomb.io/llms.txt