Cribl signals: overhead and sizing

Use this guide to estimate metric output and assess the additional work on Cribl workers. The Cribl signals recipe aggregates events as they pass through Cribl; it does not query the original Splunk log dataset. For daily disk growth, replication and retention, see Splunk storage estimation.

Download the sizing assumptions as Markdown.

What drives the overhead

There are three main contributors:

  • Per-event processing: preserve identity and timestamps and update the aggregation for every eligible event. Original log volume alone is not enough to estimate CPU: the same byte volume can represent very different event rates.

  • Aggregation state: retain active host/index/sourcetype groups in each contributing worker process. Additional dimensions and longer windows can increase the number of groups held at once.

  • Summary delivery: serialize, buffer and send the generated metrics. Longer windows can reduce this work, but early flushes and process distribution affect the actual output rate.

The recipe’s 64 MB aggregation limit is per worker process, not a total worker memory cap. Cloning, other functions, serialization and destination queues also consume memory. Increasing the window is therefore not a guarantee of lower CPU or peak memory. Use the formulas below for volume planning and the measurement procedure to determine resource requirements.

Estimate the summary volume

Use the following variables:

  • H: hosts in scope.

  • K: mean index/sourcetype feeds per host.

  • A: fraction of feeds active within an aggregation window.

  • F: mean number of processes contributing a summary for each active feed/window. This is not the total worker count.

  • T: aggregation window in seconds.

  • M: numeric metrics per summary; this recipe uses three.

Logical feeds                  = H × K
Summary equivalents / second  ≈ H × K × A × F / T
Numeric measurements / second ≈ H × K × A × F × M / T
Numeric measurements / day    ≈ measurements / second × 86,400

This assumes one flush per active process/feed/window. Early flushes, routing duplicates, retries and changes to dimensions can alter the result. Worker load balancing and connection distribution can change F. Increasing the window also changes which feeds are active, so A need not remain constant.

The following are arithmetic scenarios, not measured deployment estimates. They assume 200,000 hosts, every feed active (A=1), one contributing process (F=1), and three metrics (M=3):

Volume scenarios

Feeds / host

Logical feeds

Window

Summaries / s

Measurements / s

Measurements / day

1

200,000

60s

3,333

10,000

864 million

2

400,000

60s

6,667

20,000

1.728 billion

5

1,000,000

60s

16,667

50,000

4.320 billion

1

200,000

300s

667

2,000

172.8 million

2

400,000

300s

1,333

4,000

345.6 million

5

1,000,000

300s

3,333

10,000

864 million

If two processes usually see each feed during a window, double these values. Sparse feeds reduce them. A longer window reduces nominal output, but every eligible original event still needs processing, freshness becomes visible later, and more distinct feeds may remain in aggregation state.

Five-minute windows: estimated savings

The guide keeps a 60-second default. A five-minute window is an optional volume reduction for estates that can tolerate slower updates. Follow Choose the aggregation window: 60 seconds or five minutes for the configuration steps and monitoring tradeoffs.

Holding active feed count and process fan-out constant, with one flush per active process/feed/window:

Five-minute points = 60-second points × 60 / 300
                   = 60-second points × 0.2
Reduction          = 1 - 0.2 = 80%

For the 200,000-host, two-feed scenario above (A=1, F=1, M=3):

Estimated daily output, not a measured production result

Measurement

60 seconds

Five minutes

Daily reduction

Summary equivalents

576 million

115.2 million

460.8 million

Metric points (numeric measurements)

1.728 billion

345.6 million

1.3824 billion

These are fewer samples of the same metrics and dimensions, not 80% fewer distinct time series or monitored entities. An otherwise identical continuously active estate with 5,000 hosts and two feeds per host would go from 43.2 million to 8.64 million metric points/day. This smaller scenario is also arithmetic, not a measured deployment result.

The 80% ratio is most useful for continuously active feeds. A feed that sends only one event every ten minutes may still need one summary per event with either window, so it may see little reduction. More processes can contribute to a feed over five minutes, and early flushes can create additional summaries. Re-estimate A and F for each window instead of applying 80% to every deployment.

Validate with settled, equal-duration measurements covering several complete five-minute windows and comparable input. Exclude the configuration-change interval and allow delivery to complete. Compare the number of count-metric samples, not just the sum of their values: fewer samples should still represent the same input event population over equivalent, fully delivered windows. Check peak memory and flush-time queue pressure as well as average CPU.

Do not infer CPU from host count

For illustration, 100 TB/day of original logs corresponds to approximately 1.16 million events/s at an assumed mean event size of 1,000 bytes (decimal units), distributed across the estate’s workers. Per-event work and active grouping state must both be measured; scaling host count alone does not predict them.

For the two-feed, 60-second scenario, there are 576 million summary equivalents per day. An assumed effective serialized size of 300–1,000 bytes per summary, including all three measurements and dimensions, gives 172.8–576 GB/day of additional metadata. These are sensitivity inputs, not measured encoding sizes. They exclude protocol overhead and do not predict compression, license usage, index storage, replication or retention cost.

Measure actual destination bytes and retained Splunk footprint before costing the design. See Cribl signals: Splunk storage estimation for the storage formulas. Additional provisioned worker capacity may be unnecessary while there is sufficient headroom, but the branch still consumes resources. Compare that cost with the measured reduction in distributed monitoring searches using your own commercial terms.

Repeat the measurements in your environment

Measure a settled window with the activity branch off and another with it on, using the same input replay where possible. Keep original log processing active. Repeat the cycle; compare the busiest process as well as the group average.

This search measures a 15-minute activity window. Use explicit time bounds and the corresponding duration when comparing runs:

| mstats
    count(telemetry.activity.events.count) AS summaries
    sum(telemetry.activity.events.count) AS represented_events
  WHERE index=cribl_metrics earliest=-15m latest=now
  BY host data_index data_sourcetype
| stats dc(host) AS hosts count AS feeds
        sum(summaries) AS summaries sum(represented_events) AS represented_events
| eval measurements=summaries*3,
       summaries_per_second=round(summaries/900,2),
       measurements_per_second=round(measurements/900,2),
       represented_events_per_second=round(represented_events/900,2)

For worker resources, the following example uses Cribl internal metric names and dimensions. Adapt them to your exporter and preserve both node identity and process identity when comparing several workers:

| mstats
    avg(cribl.logstream.system.cpu_perc) AS cpu_pct
    count(cribl.logstream.system.cpu_perc) AS cpu_samples
    avg(cribl.logstream.system.mem_heap_used) AS heap_bytes
    avg(cribl.logstream.system.mem_rss) AS rss_bytes
    max(cribl.logstream.aggregators.total_memory_usage) AS aggregation_peak_bytes
  WHERE index=cribl_metrics earliest=-15m latest=now
  BY group event_host cribl_wp
| eval heap_mib=round(heap_bytes/1048576,2),
       rss_mib=round(rss_bytes/1048576,2),
       aggregation_peak_mib=round(aggregation_peak_bytes/1048576,2)

See Cribl internal metrics for the underlying measurements. Missing dimensions can remove rows from a grouped search, so first confirm the names present in your metric index.

Before publishing a capacity estimate

  1. Vary event rate and active feed cardinality independently with representative event sizes, peak traffic and process distribution.

  2. Record CPU, memory, aggregation state, emitted bytes, destination pressure and queue behavior. Test restarts and failure of the metrics destination.

  3. Compare equivalent mstats and existing tracker searches: same subjects, time semantics and polling frequency. Record wall time, concurrency and search-resource demand. A screenshot of different searches is not a benchmark.

  4. Test silent subjects, replay, invalid/future timestamps, repeated identities across estates and failure of the signal path itself.

  5. Measure the complete TrackMe lifecycle at target entity and combination counts, including persistence and scheduling, before claiming estate capacity.

The production proposition is a smaller, centralized activity dataset with a clear observation point. Exact resource cost and estate capacity depend on the workload, worker distribution and available headroom.

Destination post-processing scope

The same output-rate arithmetic applies to destination signals, but use the eligible traffic reaching those destinations as the input population. Earlier filtering or sampling changes that population. Measurements for a limited route do not establish the overhead of observing an entire estate.

CPU work occurs for every eligible input event; memory follows active aggregation groups and windows per worker process. Destination fan-out can repeat that work and create duplicate summaries. A longer window mainly reduces output frequency; it does not remove per-event evaluation or guarantee lower peak memory.

For destination sizing, compare equal input rates, event sizes and feed cardinalities with the added functions disabled and enabled. Include existing post-processing, original-log throughput, worker CPU/heap, aggregation state, queue/backpressure, output errors and signal freshness. Report the number of observed destinations and copies per event. Keep original-log delivery within the deployment’s normal targets while increasing scope.

Volume such as TB/day alone cannot determine CPU cost or worker count: event rate, metadata cardinality, process distribution and remaining worker capacity are also required.