Cribl signals: Splunk storage estimation

Estimate the additional storage for Cribl activity metrics from the number of active feeds, how often summaries are emitted, and the effective stored size of each metric point. Then apply retention and capacity headroom. This gives a repeatable calculation that can be adjusted to your estate.

For example, 150,000 hosts with three feeds each, a five-minute window, 1.9 summaries per active feed/window and 64 stored bytes per point give approximately 47.3 GB/day across the cluster. Thirty days of retention is about 1.42 TB, or 2.03 TB provisioned at 70% target utilization. These are conditional estimates: the sections below explain each input and how to replace it with a value measured in your environment.

Important

64 bytes per point is an illustrative planning coefficient, not a fixed Splunk storage rate. It is a rounded reference value informed by limited measurements with RF=3/SF=2, rather than a validated long-term capacity rate. Use the sensitivity examples and calibrate before committing capacity. Dimension cardinality, compression, bucket lifecycle and replication affect the result. The higher examples are not guaranteed upper bounds.

This page covers locally retained metrics on a Splunk Enterprise indexer cluster. Original logs, TrackMe’s separate history indexes, compute, network, backups and Splunk licensing are outside the calculation. SmartStore requires separate object-storage and local-cache sizing. For Splunk Cloud, use the contracted retention and storage model rather than treating these disk estimates as a subscription price.

Use the setup guide for configuration and the overhead guide for Cribl worker sizing.

What is being counted?

A feed is a distinct combination of host, original index and sourcetype, plus any additional grouping dimensions. A host sending three sourcetypes to one index normally represents three feeds. Changing identity dimensions creates additional combinations.

The recipe emits three metric points per summary:

  • Event count for the aggregation.

  • Latest observation timestamp in Cribl.

  • Latest source-event timestamp observed.

A metric point is one numeric measurement. It is not an original log event, an HTTP request, or necessarily one serialized HEC event: multiple measurements can be delivered together. In particular, sum(telemetry.activity.events.count) counts represented original events, while count(telemetry.activity.events.count) counts samples of that metric.

Daily metric points and disk growth

Points/day = H × K × A × E × (86,400 / W) × M
GB/day     = Points/day × B / 1,000,000,000
Inputs to the estimate

Input

Meaning

Worked example

H

Hosts in scope

150,000

K

Mean feeds per host

3

A

Fraction of those feeds active in an aggregation window

1 (all active)

E

Mean emitted summaries per active feed/window, across all contributing processes

1.9

W

Configured aggregation window in seconds

300 (five minutes)

M

Numeric metrics per summary

3

B

Effective cluster-wide stored bytes per logical metric point, including replicas

64 (illustrative)

E=1 means one summary for each active feed in each window. A feed handled by several worker processes or flushed early can produce more than one summary. E=1.9 is an example, not a default or a limit. Measure it from the emitted samples. The overhead guide uses F for the contributing-process factor, assuming one flush per process/feed/window; E also accounts for repeated flushes.

K and A should describe the same population and time window. Avoid combining all-time distinct feeds with a short-window activity fraction from a different population. For changing workloads, calculate points separately for representative periods and add the results.

150,000 × 3 × 1 × 1.9 × (86,400 / 300) × 3
    = 738,720,000 points/day

738,720,000 × 64 / 1,000,000,000
    = 47.27808 GB/day across the cluster

Do not multiply this result by RF or SF again: B already includes the cluster’s physical copies. All GB and TB on this page are decimal (1 GB = 10^9 bytes; 1 TB = 10^12 bytes). A GiB is 2^30 bytes; convert units before combining values from dashboards or storage invoices.

Examples for 150,000 hosts

The following estimates assume a five-minute window, all feeds active, three metrics per summary and B=64 bytes/point, including the assumed RF=3/SF=2 cluster footprint. They show the effect of feed count and summary frequency.

Additional cluster disk growth (GB/day)

Feeds per host

E=1 summary/feed/window

E=1.9 summaries/feed/window

1

8.3

15.8

2

16.6

31.5

3

24.9

47.3

5

41.5

78.8

Test the estimate against different storage coefficients. For the three-feed, E=1.9 example, changing B gives:

Stored-size sensitivity

Cluster bytes per point

Additional GB/day

64

47.3

128

94.6

256

189.1

These coefficients are planning inputs, not a confidence interval. Measure B with representative dimension values, grouping cardinality and index settings. Changing RF/SF requires a new coefficient or the component model below.

Aggregation window and freshness

Holding every other input constant in the three-feed, E=1.9, B=64 example:

Effect of the configured window

Window

Additional GB/day

60 seconds

236.4

Five minutes

47.3

Ten minutes

23.6

Fifteen minutes

15.8

Moving from 60 seconds to five minutes gives one fifth as many configured windows, or 80% fewer nominal points. This is conditional on A, E and B remaining unchanged. Sparse activity, early flushes, process fan-out and bucket overhead can change the actual reduction.

Longer windows delay fresh information and can retain more active aggregation state. They do not remove per-event processing or guarantee lower Cribl memory usage. The recipe keeps 60 seconds as its default; consider five minutes when lower volume is more important than faster visibility. Ten- and fifteen-minute rows illustrate the arithmetic, not a recommended default. See Choose the aggregation window: 60 seconds or five minutes for configuration, flush settings and tracker lookback.

Replication and search factor

For a fully replicated, single-site cluster using local storage, a useful component model is:

Cluster bytes ≈ RF × R + SF × I + O
RF=3, SF=2    ≈ 3 × R + 2 × I + O

R is the compressed raw-journal size of one copy, I is the searchable index-file size of one copy, and O is other cluster-wide bucket and filesystem overhead. Apply the terms to the same data population. Searchable copies are included within RF: RF=3/SF=2 does not mean six full copies. See Splunk’s bucket-copy model.

Metrics use dedicated index structures, and floating-point compression is enabled by default. However, do not assume that replicated metrics have no raw journal: metric.stubOutRawdataJournal does not take effect for indexes with repFactor=auto in an indexer cluster. See the Splunk indexes.conf reference.

This is why multiplying JSON payload size by six, or applying a generic event index compression ratio, is not a reliable metrics-storage calculation. Prefer a measured B for the intended configuration. For multisite clusters, use the effective site replication/search policy and copy distribution.

Retention, provisioned volume and price

Retained GB    ≈ daily GB × retention days
Provisioned GB ≈ retained GB / target utilization

At 47.27808 GB/day, using an illustrative target utilization of 70%:

Additional storage for the worked example

Retention

Retained TB

Provisioned TB at 70%

30 days

1.42

2.03

90 days

4.26

6.08

180 days

8.51

12.16

Retention is an approximate steady-state calculation; bucket boundaries and workload changes affect the actual footprint. On four evenly loaded indexers, 2.03 TB is approximately 507 GB additional provisioned storage per indexer. Add existing indexes and account for imbalance, maintenance and recovery. The 70% value is an example planning choice, not a Splunk requirement or a substitute for a deployment-specific capacity margin.

If storage is billed on provisioned capacity at a price P per GB/month:

Monthly storage charge ≈ provisioned GB × P
30-day example         ≈ 2,026 × P

If billing is based on actual occupied storage, use average billable occupancy instead. Add any separate IOPS, throughput, backup or transfer charges. Daily disk growth, retained capacity and monthly cost are different quantities; this calculation is not a Splunk license-usage estimate.

Calibrate the estimate in your environment

  1. Use a dedicated metrics index or another method that isolates only the signal metrics. Keep the grouping dimensions, aggregation settings and replication policy representative of the planned deployment.

  2. Choose matching start and end times over several complete aggregation windows. Count the logical metric samples once through a normal distributed search. Do not sum independent searches of replica copies.

  3. Measure that index’s footprint across all indexer peers, including all physical copies, at both boundaries. Wait for replication to settle and use consistent units and measurement methods. Dashboard index size and filesystem allocated bytes are not interchangeable.

  4. Ensure no freezing, deletion, tier movement or unrelated data changed the measured footprint. A nearly full index or shared volume can remove old data while new data arrives, making net growth understate storage demand.

  5. Repeat over representative periods, ideally several days, covering normal bucket rollover and busy traffic. A new index’s one-time size divided by its sample count is only an initial indication, not a stable growth coefficient.

Use explicit, matching time-picker boundaries with this search (replace the index name if necessary):

| mstats
    count(telemetry.activity.events.count) AS p_count
    count(telemetry.activity.last.observed.epoch) AS p_seen
    count(telemetry.activity.last.event.epoch) AS p_event
  WHERE index=cribl_signals
| eval total_points=p_count+p_seen+p_event

Allow delivery to complete for the selected interval. Check all three counts: they should be consistent with the recipe. Unexpected differences should be investigated before using the result for sizing. The index name above selects the metrics storage index, not the original-log data_index dimension.

For intervals without data expiry or other footprint changes:

B ≈ (cluster bytes at end - cluster bytes at start) / new logical points

Measure E by dividing emitted count-metric samples by the number of active feed/window combinations over the same period, using the configured window and all grouping dimensions. Window alignment and flush timing can affect a short measurement. Recalibrate after changing dimensions, window/flush settings, worker distribution or the Splunk storage configuration.

The result is a traceable storage estimate: metric output drives daily growth, retention determines the retained dataset, and headroom and commercial terms determine provisioned capacity and price.