# Cribl signals: sizing assumptions

Companion to the [setup guide](https://docs.trackme-solutions.com/latest/cribl_logstream_signals.html),
[overhead guide](https://docs.trackme-solutions.com/latest/cribl_logstream_signals_sizing.html)
and [storage guide](https://docs.trackme-solutions.com/latest/cribl_logstream_signals_storage.html).

## Observation boundary

Signals describe activity observed in Cribl. They do not prove that Splunk
indexed the corresponding original logs. This recipe emits three numeric
metrics per summary: an event count, the latest Cribl observation timestamp,
and the latest source-event timestamp. Counts are non-cumulative.

## Volume model

H = hosts; K = feeds per host; A = fraction of feeds active per window;
E = emitted summaries per active feed/window across all contributing processes;
W = window in seconds; M = metrics per summary (three).

```
Points/day = H * K * A * E * (86400 / W) * M
GB/day     = Points/day * B / 1000000000
```

B is effective cluster-wide stored bytes per logical metric point, including
replication at the measured configuration. Do not apply replication again.
A feed includes host, original index, sourcetype and any extra grouping dimensions.
Count samples of the count metric to estimate summaries; summing its values
measures original events instead.

The overhead guide uses F for contributing processes, assuming one flush per
active process/feed/window. E additionally accounts for multiple flushes.
Longer windows can change both activity and the number of contributing processes.

## Example assumptions

For 150,000 hosts, three feeds per host, A=1, E=1.9, W=300, M=3 and an
illustrative B=64 bytes/point, the result is 738.72 million points/day and
47.28 GB/day of additional cluster storage. B=128 and B=256 give 94.56 and
189.11 GB/day. These are sensitivity scenarios, not guaranteed limits or
Splunk storage constants.

With 30 days of retention, 47.28 GB/day is about 1.42 TB retained. At an example
70% target utilization, provision about 2.03 TB, before adding other workloads
or deployment-specific recovery requirements. GB and TB here are decimal units.

## Worker resources and validation

Every eligible event still needs processing. Longer windows can reduce output
but can retain more active groups and do not guarantee lower peak memory.
Compare equivalent traffic with the monitoring branch disabled and enabled;
measure CPU, heap/RSS, aggregation state, destination queues and signal freshness.

Calibrate storage using several representative intervals with complete cluster
replication and matching metric counts and disk measurements. Avoid freezing,
bucket deletion or tier movement during the measurement. Include dimensions,
retention, replicas, load distribution and capacity headroom in your estimate.

For RF=3/SF=2, a component model is 3 * raw-journal bytes + 2 * searchable-index
bytes + other overhead, not six full copies. SmartStore and Splunk Cloud require
their own storage and commercial model. Splunk licensing, compute, network,
backups, original logs and TrackMe's separate history are outside this estimate.
