Product introduction¶
What TrackMe is, what it tracks, how it decides and how it tells you — the vocabulary and the mental model before the live product.
about 20 minutes · 16 slides. Each slide is followed by a description of what it shows — read top to bottom, this page is the walkthrough.
Overview¶
TrackMe for Splunk is presented through its key concepts: what the product is, what it tracks, how it decides and how it tells operators. The material covers TrackMe 2.4.18 and maps one to one to the Product Guide on docs.trackme-solutions.com, which remains the reference for every term used here.
The topics follow the product’s own vocabulary: entities, components and virtual tenants, then the decision maker, policies and alerting, followed by Topology Studio, the AI layer, the architecture and the typical adoption path.
Why Splunk data breaks silently¶
Four failures share one root cause. A forwarder stops, yet the index keeps ingesting from other hosts and the dead machine is invisible. A sourcetype arrives late, so detections run on a window that no longer contains the events they need. A field stops parsing, and correlation searches quietly return less, or nothing. A scheduled search skips, summary indexes and reports drift, and scheduler debt accumulates over years.
Splunk itself does not report any of this: searches simply return less until a detection misses or a dashboard goes blank. Modern Splunk estates hold thousands of indexes, tens of thousands of sourcetypes and hundreds of thousands of hosts, far beyond what can be watched or classified by hand.
Definition and the four open questions¶
TrackMe is a Splunk app that discovers, maintains and monitors the availability and quality of data, at any scale. It is a control plane on top of Splunk: raw index activity becomes per-entity health states, SLOs, stateful alerts and, optionally, LLM-generated investigations, entirely inside Splunk with no outbound SaaS.
It answers four questions classic Splunk monitoring leaves open. Is the data still flowing: data sources, hosts and metric feeds are tracked for delay, latency and volume, plus licensed volume per index. Is it still good: parsing quality and CIM compliance are measured. Are searches healthy: execution, skipping and errors are first-class entities. Who to tell, and how urgently: priority, SLA and impact scoring separate alerts from noise.
The six terms of the vocabulary¶
One sentence holds the model: TrackMe tracks entities (grouped into components, isolated inside virtual tenants); on every cycle a decision maker gives each entity a state and a score, policies classify it, and alerting turns critical states into notifications. Callouts 1 to 6 name the six terms.
An Entity is the unit monitored. A Component is its kind: DSM, DHM, MHM, VOL, FLX, FQM or WLK. A Virtual Tenant is an isolated domain with its own entities, configuration and roles. The Decision maker computes state and score once, so UI, alerts and API agree. A Policy assigns priority, tags and SLA by rule; a manual override always wins. Alerting is stateful: one incident record per problem.
Component families and their questions¶
The seven components are grouped by the question they answer. DSM (index and sourcetype feeds), DHM (host and sourcetype activity) and MHM (metric hosts) answer whether the data is arriving; they form the Splunk Feeds family and share a delay, latency and volume model. VOL, Volume Outliers, checks the licensed volume of each index with ML outliers on drops and spikes.
FQM, Field Quality Monitoring, measures parsing quality and CIM compliance per field. WLK, Workload Knowledge, tracks scheduled-search skipping, errors and late runs. FLX, Flex Objects, turns any SPL into a tracked entity. UAM, User Activity Monitoring (Beta), is a dedicated tenant type, not a component: it watches Splunk accounts and raises findings from a policy catalogue.
Tenants, groups and remote deployments¶
A virtual tenant is a self-contained, isolated instance of TrackMe inside the Splunk app. It owns its data, configuration, trackers, access control and, optionally, its own remote Splunk deployment to monitor. The Virtual Tenants page is the entry point of the product. Tenants carve an estate into domains: by company, business unit or team; by technology, with per-technology thresholds; by scope, one tenant per cluster, index family or search-head tier; and for access control.
Everything else is scoped to a tenant; there is no global entity list. Virtual Groups aggregate entities from several tenants into one read-only cross-tenant card. A tenant can monitor a remote Splunk Enterprise or Cloud over a token-rotated connection with failover, with an identical experience.
Discover, measure, decide, classify, protect, alert¶
TrackMe is not a live dashboard that queries Splunk on demand. Trackers run on a cycle and persist their results; the UI reads the last computed state, which is what lets the product scale and alert unattended.
Callouts 1 to 6 follow one cycle. DISCOVER: trackers scan the tenant’s indexes for new (index, sourcetype) pairs, hosts and searches. MEASURE: delay, latency, volume, parsing quality and execution health are refreshed. DECIDE: the decision maker computes the impact score and the state green, orange or red. CLASSIFY: policies assign priority, tags and SLA. PROTECT: logical groups, disruption grace, maintenance and acknowledgment may demote an entity to blue. ALERT: stateful alerting opens, updates and closes records and delivers notifications.
Hybrid impact scoring and the four colours¶
Each cycle, every entity receives one of four colours derived from a numeric impact score. The total_score starts from the base_score and adds impact_score_delay and impact_score_latency when those thresholds are breached, impact_score_outliers from ML outlier detection and threshold_scores for breached dynamic thresholds, each a configurable weight.
The score maps to a colour: 0 is green, no active anomaly; below 100 is orange, an early signal that does not page; 100 and above is red, and the alert fires by design. Blue means suppressed: group protection, disruption grace or maintenance protects the entity from alerting although an anomaly may be present. Weights are configurable per tenant and overridable per entity, so nothing happens without a traceable reason.
Metadata, policies and protection mechanisms¶
Classification attaches four kinds of metadata. Priority says how much an entity matters. Tags are free-form strings for filtering and grouping. Labels come from a curated, colour-coded catalogue of chips. The SLA class is platinum, gold or silver with associated thresholds. Policies assign metadata by regex, Splunk lookup or SPL search; a manual override always wins, and removing a policy auto-cleans its assignments.
Four mechanisms hold an entity in blue instead of red. Logical group protection judges entities collectively and demotes struggling members when enough are green. Disruption grace holds a new anomaly until it has persisted for a minimum time. Maintenance mode covers a window; acknowledgment records that someone accepted the problem.
Incident lifecycle and delivery channels¶
Stateful alerting keeps one incident record per problem. INCIDENT OPENED fires when the entity first turns into an alerting state, red or, by opt-in, orange; INCIDENT UPDATED follows while the problem persists; INCIDENT CLOSED fires when the entity recovers. Each transition writes a stateful event to trackme:stateful_alerts, can send a threaded email (Message-ID on open, In-Reply-To afterwards) with an optional AI report, and can run an active command through commands_opened, commands_updated or commands_closed.
Four delivery channels exist: Email with embedded 24-hour charts and an optional AI-written status report; Notable events for Enterprise Security or a custom correlation workflow; Ingest, an event written to a Splunk index; and Commands, a saved search or active command towards SOAR, XSOAR or ServiceNow.
Node types, sharing and Topology Alerts¶
Topology Studio builds live dependency maps from real TrackMe entities, with status updating in real time. An Entity node is one monitored entity with its live state colour and optional KPI. An Aggregate node is a live roll-up of many entities, showing the worst state and per-state counts. A View node folds another saved view into one health node. A Label is a free-text annotation for zones, tiers or flows.
Views are global objects shared under RBAC with separate read and write permissions. Topology Alerts correlate the health of the whole map and notify on transitions. Topology Studio is available in every edition and complements the per-component topology graphs and Virtual Groups.
Four AI levels and their guardrails¶
The AI layer helps operators understand and resolve what TrackMe detects, grounded in live state. Level 1, the AI Assistant, is a context-aware chat drawer on every page that can write AI status reports into alert emails. Level 2, AI Advisors and Concierge: six specialists that inspect and, with consent, act on ML models, feed lifecycle, FLX thresholds, field quality, component health and UAM findings. Level 3, AI Agents automation, runs them unattended nightly. Level 4, AI Routines, evaluates standing instructions in natural language.
The layer is opt-in and inert until an administrator configures a provider. Every claim traces to a tool call, inspect mode is the default, and every write is prefixed [AI Agent] in the audit trail.
Data plane, control plane, presentation¶
TrackMe is built in three layers, all inside Splunk. The Data plane is the trackers and a shared decision maker that computes each entity’s state and score. The Control plane, a REST API and a per-tenant KV Store, holds all configuration and current state. The Presentation layer is the React interface, Virtual Tenants, Tenant Home, Topology Studio and dashboards, plus the optional AI Assistant. The UI reads state; it does not compute it.
TrackMe runs on Splunk Enterprise or Cloud with no outbound SaaS, and is in production above 100 TB/day. Editions are Foundation, Enterprise and Unlimited; a new install starts a 90-day trial, and an expired licence first switches to read-only while monitoring keeps running.
Six adoption phases¶
Most deployments walk the same path; many are in several phases at once. Phase 1, Data sources with DSM, is where every deployment starts. Phase 2, Hosts and endpoints with DHM and MHM, detects a single machine going silent, usually in a dedicated tenant. Phase 3, Volume outliers, tracks the licensed volume of every index with VOL.
Phase 4, Scheduler health with WLK, surfaces the scheduler debt large estates accumulate. Phase 5, Expansion with FLX, FQM, UAM and Topology, covers Splunk infrastructure, field quality, platform activity and topology maps. Phase 6, Steady state, is a continuous practice. Common mistakes: mixing DHM into the feeds tenant, enabling ML on every component, and enabling every FLX template at once.
Where we are¶
Callouts 1 to 6 recap the model. TrackMe is a Splunk app that discovers, maintains and monitors the availability and quality of data. Seven components, DSM, DHM, MHM, VOL, FLX, FQM and WLK, share one entity model; UAM (Beta) is its own tenant type. Virtual Tenants are the unit of isolation and scale, with their own entities, configuration and roles. The cycle runs discover, measure, decide, classify, protect and alert each tracker run.
Green, orange, red and blue derive from one impact score. Policies classify, protection mechanisms hold entities in blue, and stateful alerting turns critical states into one incident per problem; Topology Studio and the AI layer sit alongside. Reference: the Product Guide on docs.trackme-solutions.com.