Volume Outliers — in depth

Tip

This is the in-depth reference for the Volume Outliers (splk-vol) component. For the readable overview of what VOL is and when to use it, start with VOL — Volume Outliers; for a guided, end-to-end walkthrough on a live deployment, read the Volume Outliers white paper. This page covers the collector, license days and time zones, the backfill, the outliers and inactivity policies, the projection, the Global license usage view, alerting, the REST endpoints, the metrics and troubleshooting.

The entity model

  • An entity is one Splunk index: object is the index name, object_category is splk-vol, and the entity’s group is its license pool (a single-pool deployment has one group; on Splunk Cloud, where a pool has no limit, the group is default).

  • The state is green by default: collection never turns an index orange or red on its own. The state comes from ML outliers (a lower-bound breach is a drop, an upper-bound breach a spike) and from inactivity, through the same hybrid impact scoring as every component (Entity State & Scoring). The tenant impact keys are impact_score_vol_outliers and impact_score_vol_inactive (default 100: an inactive index turns red).

  • Per index the analyst chooses the detection direction — Detect drops, Detect spikes, or both — with the inline Detect toggle of the VOL table, in bulk edit, or in the entity’s Volume & cost tab. The direction is stored on the index’s outliers model (alert_lower_breached / alert_upper_breached); the tenant default is vol_default_direction (both).

  • Everything generic applies: priority, tags, labels, SLA, logical groups, disruption queue, acknowledgments, notes, maintenance, monitoring time, enable / disable, deletion. An index deleted temporarily keeps its bucket history when it had licensed volume in its last 24 hours (it comes back with its true rolling volume); a permanent deletion drops it and the index is never collected again.

The collector

One report per tenant, trackme_vol_tracker_tenant_<tenant_id>, runs every 5 minutes through the standard tracker executor (execution summary, licensing checks and the scheduled-time anchor come with it). Each run:

  1. Loads the checkpoint of the tenant — the cursor (the end of the last finalized 5-minute bucket), the validated history start, the license manager’s UTC offset and the deployment type.

  2. Computes the window [cursor, floor(now − settle)), aligned on 5-minute boundaries. The settle margin (vol_settle_seconds, default 300 s) leaves the license manager the time to write every Usage event of a bucket before the bucket is read (on a large deployment those events land within seconds; the margin covers the worst case). A run catches up at most 6 hours at a time; a gap above 24 hours re-seeds the history (see Validated history below).

  3. Runs one bounded search on the license usage log, local or on the tenant’s remote account:

    search (index=_internal source="*/license_usage.log") type=Usage idx=* earliest=<cursor> latest=<end>
    | bin _time span=300s
    | stats sum(b) as b, latest(poolsz) as poolsz by _time, idx, pool
    

    plus, once a day (and on every run around the license manager’s midnight), a cheap indexed probe on the RolloverSummary event that returns the license manager’s UTC offset and the deployment type.

  4. Updates the ring of each index — the last 288 buckets (24 hours) kept in the tenant’s KV store: appends the new buckets (a bucket with no event is 0), recomputes the rolling sums at each bucket end, accumulates the current license day and closes any license day that ended.

  5. Writes the metrics with explicit bucket timestamps, then advances the cursor — the cursor moves only once the metrics and the rings are persisted, so a run that dies mid-way is replayed, never skipped.

  6. Emits one entity row per known index, carrying the latest rolling values, the daily figures (yesterday, month to date, 30-day average, trend, projection), the pool usage on Splunk Enterprise, the detection direction and the effective inactivity policy.

Note

Why a ring instead of querying the metrics index. The rolling volume must not depend on the metrics pipeline’s own indexing latency: a blocked queue would read as a drop and raise false lower-bound outliers. The rolling values are computed from the buckets TrackMe holds and then written as metrics.

Integrity rules

  • Validated history. Every ring and the tenant checkpoint carry the time from which the history is continuous. A rolling KPI is emitted only once its whole window lies inside the validated history, and a license day is reported only when it was observed from its start — the first run’s 24-hour seed writes the 5-minute volumes for every bucket, but the 24-hour trend only from the last seed bucket. No ramp-up or gap ever reaches the series the models train on.

  • Search errors fail the run. An error from the usage search — or a warning saying the results are incomplete (a peer down, a search process ended prematurely) — stops the run before anything is written; the cursor does not advance and the missing volume is never recorded as a drop.

  • One collecting run at a time. A lock protects the collection; an overlapping run (for instance run the tracker now during the schedule) rebuilds the entity rows from the stored rings without writing any metric point. A run that meets the lock of a tenant being seeded waits for it (up to 4 minutes) and then collects normally.

  • A new source re-seeds. Changing the target environment (vol_account) or the license search constraint discards the rings and the checkpoint and re-seeds 24 hours from the new source, emitting only the points after the previous source’s cursor — a series never holds two points for the same bucket. The previous source’s indexes stop being updated, turn inactive and can be deleted.

  • Fail-closed reads. The rings, the per-index directions and the permanently deleted indexes are read without a fallback: a failed read fails the run instead of resetting a direction or resurrecting a deleted index.

  • Squashed rows keep the index. When the license manager squashes Usage rows (squash_threshold), it drops the host and source, never idx — squashing never reads as a drop.

License days and time zones

Splunk measures the license on license days that roll at the license manager’s local midnight, whatever the time zone of your search heads or of your Splunk user. VOL follows the license manager:

  • The daily probe reads the license manager’s UTC offset from the RolloverSummary event (its raw timestamp carries the offset, and its time is the true day boundary). On a deployment younger than a license day no rollover has been logged yet: the offset is then read from the raw timestamp of the latest Usage event (every line of the license usage log carries it) and the Manage: volume collection screen says so. UTC is assumed only when there is no license event at all, and the open day is voided when the real offset turns out to differ.

  • The probe runs on every run between 22:00 and 03:00 license-manager time, so a daylight-saving transition is caught within the run that follows it. The buckets between the old and the new day boundary move to the day they belong to, and a corrected closed day is re-emitted (a fall-back day lasts 25 hours, a spring-forward day 23 hours).

  • The daily_volume_mb metric of a license day is stamped in the middle of that day, so a chart binned by day places it on the right date whatever the viewer’s time zone.

The daily figures (yesterday, month to date, the trend and the projection) refer to license days, and the Volume & cost tab says which license date “today” is for the collector.

Backfill

At creation the tenant requests a backfill of its history (vol_backfill_days, default 30, 0 to 90) so the outliers models train immediately and the daily history, the trend and the month projection are complete from the first day. The wizard caps the depth to the retention of the license usage log it can reach and shows the estimated cost in metric points.

  • The backfill is a separate report, trackme_vol_backfill_tracker_tenant_<tenant_id>, scheduled hourly while a plan is open and un-scheduled once complete. It walks forward in time (every rolling value needs the 24 hours before it) in day-sized chunks, each chunk replaying exactly the live collection with historical timestamps, and runs for at most 50 minutes per slot.

  • The backfilled range ends where the live collection started: the live tracker and the backfill never overlap and no point is written twice. The first backfilled day is aligned on a license-day start, so it is complete; every complete license day of the range emits its daily_volume_mb.

  • Daylight-saving transitions inside the range are replayed with the offset in effect at the time, from the RolloverSummary history.

  • On completion the backfilled days are merged into each index’s live history (the entity figures reflect them at once), the seam day around the live start is completed, and a training of every VOL model is requested, so outliers are usable without waiting for the next training cycle.

Progress is visible in three places — the tenant card on the Virtual Tenants page (Backfilling 42 %), the vignette in the Tenant Home header band, and the Manage: volume collection screen (live cursor, validated history, license-manager offset, deployment type, the backfill progress with pause / resume / cancel / restart, and run the tracker now). From the creation until the backfill job’s first hourly run, the three read Backfill pending with the requested depth and the job’s next run.

Note

A restart re-plans only the history before the earliest point already written: it extends the history backward and never rewrites it. A change of the target environment or of the license search constraint while a plan is open cancels the plan (the new source is re-seeded first); restart it once the live collection runs on the new source.

The license search constraint (advanced)

Every license search of the tracker, the backfill and the probes is formed from one tenant-level constraint, vol_license_search. Its default targets the license usage log of the deployment the tenant collects from:

index=_internal source="*/license_usage.log"

You normally never change it. The Advanced section of the wizard’s Volume step (and the Configure tenant form afterwards) lets you point the tenant at another data set holding events in the license_usage.log format — for instance a copy of the license usage events forwarded to a dedicated index, or a deployment where _internal is not reachable under that name. The constraint is a plain base search on indexed fields:

  • no pipe, subsearch (square brackets), macro (backtick), line break, earliest or latest — the tracker sets the time range;

  • it must name an index as a real search term (index= inside a quoted value does not count), and not every index: a bare wildcard (index=*) is refused, index=lic* is fine;

  • every OR alternative must be index-scoped (index=a OR index=b is fine, index=a OR source=x is refused) and parentheses must balance.

The wizard validates the constraint as you type, and it is validated again at every save.

Important

The data behind a custom constraint must have the license_usage.log key=value fields extracted at search time — type, idx, b, pool, poolsz — since the tracker filters on type=Usage idx=* and sums b. The real license_usage.log in _internal (sourcetype=splunkd) is extracted out of the box; a copy stored under a sourcetype with KV_MODE = none (stash, for instance) needs a props.conf stanza with KV_MODE = auto for its source, or the tracker reads no usage row at all.

Changing the constraint on a running tenant re-seeds the collection from the new source (see Integrity rules above) and cancels an open backfill plan.

Outliers

VOL rides TrackMe’s native outliers engine (Machine Learning, Outlier detection — in depth) with component-specific defaults:

  • Every new index gets one model per default KPI — vol_outliers_kpis, default last_24h_volume_mb, any subset of the four rolling KPIs. The rolling 24-hour volume follows the daily pattern and is the proven choice; the shorter windows (12 h, 4 h, 60 min) react faster to a sudden drop or spike and are noisier. One model per index is the deliberate default: several models per index multiply the noise. The models use the day-of-week × hour-of-day time factor over a 90-day calculation period.

  • The default KPIs are set in the wizard and changed later in Tenant Home (Manage: Outliers default KPIs), with an opt-in apply to the existing indexes that adds the models of newly selected KPIs and removes the models of deselected rolling KPIs — models of kept KPIs and of non-rolling metrics are untouched. A model can also be added on, or switched to, any of the four KPIs from the entity’s Outliers tab.

  • Direction — a drop is a lower-bound breach, a spike an upper-bound breach; the per-index detection direction enables one or both bounds on the model.

  • Scope — VOL has its own scope: vol_mloutliers_priority_filter (default: every priority — volume outliers are the point of the component) and vol_mloutliers_filter_expression (the TrackMe filter DSL on object, priority, tags, labels and pool). The tenant-wide mloutliers_priority_filter and mloutliers_filter_expression do not apply to VOL entities. The Manage: ML Outliers scope modal edits the VOL pair on a VOL tenant, and the table’s Outliers column shows out of scope with the reason for an excluded index.

  • Minimum history — the wizard exposes the tenant override of splk_outliers_min_days_history (15 days on a VOL tenant by default, against a system default of 30) and tells you whether the requested backfill covers it: with 30 days of backfill, the models reach normal confidence on day one.

  • ML is always on for VOL. A tenant where ML is off for the other components because it only runs DHM / MHM keeps ML on for VOL alone (mloutliers_allowlist=vol).

  • The AI ML Advisor — interactive and its automated batch — covers VOL models like any other (AI ML Advisor).

Inactivity policy

An index whose last licensed volume is older than the threshold in force turns red (anomaly_reason=inactive, impact_score_vol_inactive). Two policy shapes exist:

  • static — one threshold (vol_max_sec_inactive, default 7 days);

  • variable — slots by day of week and hour of day, the DSM / DHM variable-delay model: {slot_name, days (0 = Monday .. 6), hours (0-23), max_sec_inactive}; the first matching slot applies, the default (vol_inactive_default) outside every slot.

The tenant policy (vol_inactive_policy, variable by default: 1 day during working hours, Monday to Friday 08:00–19:59, and 7 days for nights and week-ends) applies live to every index without its own override. A per-index override — static or variable — is set from the entity menu (Modify → Inactivity threshold), from the bulk edit category Inactivity Threshold, or by clicking the Inactivity column, and can be sent back to the tenant policy. A threshold of 0 disables the check (the whole static policy, or one slot); otherwise it runs from 300 seconds to 90 days. Durations accept seconds, a unit suffix (30m, 12h, 1d, 1w) or [D+]HH:MM:SS.

Note

The slot clock is splunkd’s system time zone — the frame the editor’s hour grid is drawn in (the editor shows the server time and converts your browser hours). It is not the time zone of the user running the tracker or viewing the page, so the scheduled tracker, a Tenant Home refresh and an inactivity save all pick the same slot.

The red status message names the slot in force (and says set for this index when an override applies); the table’s compact Inactivity column shows the threshold in force, with the slot and the source in its tooltip.

Trend and projection

The collector computes, per index and on every run, from the closed license days it holds:

  • yesterday_volume_mb, mtd_volume_mb (month to date), avg_daily_30d_mb;

  • trend_30d_pct — the least-squares slope over the last 30 closed days, as a percentage of their mean (from 7 days of history);

  • projected_month_mb — the month to date plus, for each remaining day, its expected rate: the average of the same weekday over the last 28 license days once two samples exist, else the mean of the last 7 days. The open day counts for the larger of its volume so far and its expected rate. No projection before 7 closed license days: a day or two of history gives no sensible monthly estimate. A day missing inside the index’s history (a collector outage past the catch-up) is estimated at its rate; the days before the index’s first recorded day are not.

The Volume & cost tab draws the daily licensed volume of the last 60 days with the per-day projection to the end of this month or of next month (a switch remembered per viewer), scaled to the collector’s figure so the chart and the card agree; Open in search carries the same rule in SPL. There is deliberately no predict: a weekday-weighted rate is what the license bill follows.

Global license usage (VOL)

A Tenant Home tab next to the VOL component tab: the license consumption of the whole environment, read-only. It reads what the tenant already keeps (the metrics and the entity figures) — no monitoring, no alerting, no search on the license usage log — so it works the same for a remote deployment and adds no cost.

  • Period — 30, 60 or 90 closed license days (the day boundaries follow the license manager’s time zone). A backfill still filling the history in, or a history shorter than the period, is stated in the view.

  • Key figures — licensed yesterday, daily average, peak day, the 30-day trend of the environment, month to date, month projection (the sum of the per-index projections), the share of the top 5 consumers, and on Splunk Enterprise the quota used yesterday.

  • Insights in plain words — days over quota and pools close to it, the trend, the largest increase and decrease, indexes that stopped consuming, new consumers, the concentration, the weekly pattern and the licensing model.

  • Top consumers per day — the stacked daily volume of the top consumers (a slider from 5 to 100, remembered per viewer; the rest as All others), with the daily quota as an overlay line on Splunk Enterprise.

  • The whole environment per day — the daily total with the same projection controls as the entity chart and the daily quota line when the license is metered.

  • Consumers — a sortable, filterable table (total, average, peak, days, last 7 closed days against the 7 before, share, pool, month projection, 30-day trend); a click opens the entity.

  • Movers — the week-over-week increases and decreases. They need the recorded history to cover both weeks (14 recorded days): on a young tenant the card says so instead of calling every index new.

  • Pools — one card per license pool with its usage meters (yesterday, average, peak day), the days over quota and the headroom. On Splunk Cloud the view says the deployment has no pool usage.

Every chart draws its volumes in one unit — MB, GB or TB — picked from the chart’s largest value; the tables and tiles follow the same rule (GB from 1024 MB, TB from 1024 GB).

Alerting

VOL entities flow through stateful alerting like every component. What is VOL-specific:

  • The Create Alert wizard offers the SPLK-VOL technical alert type (a filter on object_category=splk-vol); the outliers trigger stays on by default, the outliers being the VOL state. VOL is selectable in the stateful email component filter and in the SLA breaches alert.

  • Stateful alert emails embed the rolling volume charts (24 h / 12 h / 4 h / 60 min over the alert’s chart window), the outliers chart of each model, and a daily licensed volume bar chart over the last 30 license days.

  • AI status reports run the ML outliers investigation for the index.

  • Auto-acknowledgment, notable events, the free-style REST call, maintenance mode, CMDB and alert enrichment (splk_general_vol_cmdb_search, splk_general_vol_alert_enrichment_search) and the drilldown (TenantHome?component=vol&keyid=…) need nothing special.

See Alerting — in depth for the alert types and the email anatomy.

REST endpoints

Resource group splk_vol (browse it with API & tooling → TrackMe REST API Reference, every endpoint supports describe=true):

Endpoint

Purpose

POST /trackme/v2/splk_vol/vol_get_table

The VOL entity table of a tenant.

POST /trackme/v2/splk_vol/vol_entity_info

One index: its figures, direction, inactivity policy in force and the entity searches.

POST /trackme/v2/splk_vol/vol_get_state

The collector state (cursor, validated history, license-manager offset, deployment type) and the backfill progress.

POST /trackme/v2/splk_vol/vol_license_overview

The Global license usage context and searches (period_days 30 / 60 / 90, top_n 1 to 100).

POST /trackme/v2/splk_vol/write/vol_set_direction

The detection direction of one or more indexes (object_list or keys_list).

POST /trackme/v2/splk_vol/write/vol_set_inactivity

The inactivity policy of one or more indexes: policy tenant / static / variable with max_sec_inactive or slots.

POST /trackme/v2/splk_vol/write/vol_backfill_control

action pause / resume / cancel / restart, optional run_now.

POST /trackme/v2/splk_vol/write/vol_bulk_edit, vol_update_priority, vol_monitoring, vol_delete, vol_update_wdays, vol_update_hours_ranges, vol_update_monitoring_time, vol_update_manual_tags, vol_update_sla_class

The generic entity operations, as for every component.

POST /trackme/v2/vtenants/admin/update_tenant_vol_inactivity_policy

The tenant inactivity policy (dispatches the tracker so the indexes apply it within a minute); read it in the vol block of GET vtenants/tenant_default_delay_config.

POST /trackme/v2/vtenants/admin/update_tenant_vol_outliers_kpis

The default outliers KPIs, with the opt-in reconcile of the existing indexes.

The tenant is created with tenant_vol_enabled (see Creating a tenant); every vol_* option below can be set in the same call.

Tenant options

Stored in the tenant’s vtenant_account and editable from Configure tenant:

Option

Default

Meaning

vol_account

local

This Splunk, or the remote account the license usage is collected from.

vol_license_search

index=_internal source="*/license_usage.log"

The license search constraint (advanced, see above).

vol_settle_seconds

300

Settle margin (60–3600 s) before a 5-minute bucket is read.

vol_backfill_days

30

Days of history replayed at creation (0 disables, max 90).

vol_default_direction

both

Default detection direction of new indexes: both / lower (drops) / upper (spikes).

vol_outliers_kpis

last_24h_volume_mb

The rolling KPIs every new index gets a model on.

vol_mloutliers_priority_filter

every priority

Priorities in scope for outliers.

vol_mloutliers_filter_expression

empty

Filter expression restricting the indexes in scope.

vol_inactive_policy

variable

The tenant inactivity policy shape.

vol_max_sec_inactive

604800

Static policy threshold (7 days).

vol_inactive_slots

working hours

Variable policy slots (Mon–Fri 08:00–19:59 at 1 day).

vol_inactive_default

604800

Variable policy threshold outside every slot (7 days).

impact_score_vol_inactive

100

Impact score of an inactive index.

Metrics

Written to the tenant’s metric index under trackme.splk.vol.* (object_category=splk-vol, dimensions tenant_id, object, object_id):

Metric

Meaning

Unit

trackme.splk.vol.volume_5m_mb

Licensed volume of the 5-minute bucket (the base series; the hourly profile is mstats sum(…) span=1h).

MB

trackme.splk.vol.last_60m_volume_mb, last_4h_volume_mb, last_12h_volume_mb, last_24h_volume_mb

Rolling volume ending at the bucket end — the outliers inputs.

MB

trackme.splk.vol.last_24h_pool_pct_used

Share of the license pool over the last 24 hours (Splunk Enterprise only).

percent

trackme.splk.vol.daily_volume_mb

Licensed volume of a closed license day, stamped mid-day.

MB

trackme.splk.vol.status

Decision-maker status.

status code

Every point costs 150 bytes of metric-index licensing: about 288 buckets × 5 measures per index per day live (≈ 216 KB per index per day); the backfill writes the same series over the requested depth, and the wizard shows that estimate before you confirm.

The entity searches (Open in Search on every chart) read those metrics:

| mstats latest(trackme.splk.vol.daily_volume_mb) as daily_volume_mb
    where index=<metric_index> tenant_id="<tenant>" object="<index>" span=1d
| timechart first(daily_volume_mb) as daily_volume_mb span=1d
| eval daily_volume_gb=round(daily_volume_mb/1024, 3)

Troubleshooting

Symptom

What to check

The tracker succeeds but the tenant has no entity

The license search returns no Usage row: on a custom constraint, the type / idx / b fields are not extracted (add KV_MODE = auto for the source); on a remote account, the account’s roles cannot read _internal on the remote deployment.

The daily chart shows no bar, or a bar a day early

No closed license day yet (a day sits mid-day; the entity overview chart carries its own 30-day window) — or your Splunk user’s time zone differs from the license manager’s: the daily point is placed on its license date, check the user preferences.

Every index turned red at once

Inactivity in force: the collection stopped (a failed usage search, the tracker disabled, a remote account failing) — the Manage: volume collection screen shows the cursor; the execution summary shows the run errors. Fix the source, the next run catches up 6 hours at a time.

No projection on the Volume & cost tab

Fewer than 7 closed license days, or the backfill has not merged its days yet (wait for the plan to complete).

The movers card says it needs 14 recorded days

Expected on a young tenant: the week-over-week comparison needs both weeks recorded.

A backfill plan is cancelled (live_history_reset)

The live collection re-seeded (a gap above 24 hours, a source change) while the plan was open; restart the backfill from the Manage: volume collection screen once the live cursor advances again.

See also