DataKit - Data collector

DataKit is an open-source, all-in-one data integration Agent that supports all operating systems (Linux/Windows/macOS), offering comprehensive data collection capabilities across hosts, containers, middleware, tracing, logging, and security inspections.
Collect hosts, containers, Kubernetes, middleware, databases, networks, logs, traces, and custom metrics Field extraction, type conversion, filtering, and desensitization are completed through Pipeline before inbound storage Compatible with ecosystem data sources such as OpenTelemetry, Prometheus, Telegraf, Fluentd, and StatsD

Data Collection Agent

What is DataKit?

DataKit is a Guance data collection Agent that can collect metrics, logs, traces, objects, events, and custom data from hosts, containers, Kubernetes, cloud hosts, and edge environments, and access data into Guance workspaces through pipelines, unified tags, and open protocols.

DataKit features Functional features
Multi-dimensional data collection

Multi-dimensional data collection

Supports collecting data such as Metrics, Logs, and Traces from various infrastructures and technology stacks, and processing these data in a structured manner.

Kubernetes cloud-native technology support

Kubernetes cloud-native technology support

Based on a cloud-native microservices architecture, it covers comprehensive data collection within the Kubernetes ecosystem.

Observable text processing pipeline

Observable text processing pipeline

Built-in easy-to-use data extraction and processing engine, Pipeline, is used to extract unstructured data, making querying and statistics convenient.

Better tech stack support than Telegraf

Better tech stack support than Telegraf

Compared to Telegraf, which can only collect time-series data, DataKit covers a much broader range of data acquisition types, with simpler data collection configurations and better data quality.

Flexible deployment, simple and easy to use

Flexible deployment, simple and easy to use

DataKit is deployed and installed with one click, with hundreds of built-in data integrations, supporting automatic basic data collection during installation, ready to use out of the box.

Core Capabilities

From collection and processing to unified association

DataKit is not just a single collector. It handles data access, preprocessing, tag specification, ecosystem compatibility, and large-scale deployment and maintenance, bringing observable data into a unified analytical context.

Comprehensive data collection

Deploy DataKit across hosts, containers, Kubernetes, databases, middleware, and cloud resources, unifying collection metrics, logs, links, objects, and events, reducing the caliber differences caused by multiple collectors coexisting.

Pipeline data processing

Through Pipeline, log and text data are parsed, field extracted, typed, filtered, desensitized, and tag complete, making subsequent retrieval, alerting, and aggregation analysis more stable.

Open ecosystem compatibility

Supports data access from open ecosystems such as OpenTelemetry, Prometheus, Telegraf, Fluentd, and StatsD, helping existing collection systems smoothly enter Guance unified analysis.

Cloud-native deployment and governance

Supports deployment on Linux, Windows, macOS, Docker, Kubernetes, and cloud hosting, and can reduce large-scale maintenance costs through election, GitOps, DataKit API, and command-line tools.

DataKit available platform Available platforms
Linux platform
macOS platform
Windows platform
Kubernetes platform
Docker platform
Cloud hosting platform
Container platform

Workflow

Recommended access methods

01

Deploy DataKit first

Deploy DataKit on hosts, containers, Kubernetes, or cloud hosts, prioritizing access to infrastructure, critical applications, logs, and alert-related data.

02

Standardize tags and pipelines

Unified service name, environment, version, region, team, and business tags, and use Pipeline to handle log fields, anonymization, and data quality.

03

Enter Guance Unified Analysis View

Metrics, logs, links, RUM, events, and alerts are linked to the same object, forming a troubleshooting chain from problem phenomena to root cause location.

DataKit ecosystem support Ecosystem support
DataFlux Func programmable data processing development platform

DataFlux Func programmable data processing development platform

Supports receiving log data collected by Fluentd

Supports receiving log data collected by Fluentd

Supports receiving metrics collected by Telegraf

Supports receiving metrics collected by Telegraf

Scheck cloud-native programmable security inspection plugin

Scheck cloud-native programmable security inspection plugin

Supports collecting and receiving Prometheus system metrics

Supports collecting and receiving Prometheus system metrics

Supports receiving Statsd metrics

Supports receiving Statsd metrics

FAQ

Frequently asked questions

What is DataKit?

DataKit is an Guance open-source, all-in-one data collection Agent used to collect data from hosts, containers, Kubernetes, logs, links, databases, middleware, networks, cloud resources, and custom business data.

Can DataKit be used together with OpenTelemetry, Prometheus, and Telegraf?

Yes, you can. DataKit can directly collect data or receive ecosystem data from OpenTelemetry, Prometheus, Telegraf, Fluentd, StatsD, and others, helping teams integrate existing collection links into Guance unified analysis view.

How does DataKit help logs and metrics perform correlation analysis?

During the collection phase, DataKit completes unified tags such as service name, environment, host, container, Pod, and version, and processes unstructured logs through pipelines, allowing metrics, logs, links, and alerts to be associated around the same object.

If you already have custom systems or internal metrics, can you access them via DataKit?

Yes, you can. DataKit supports custom data collection, PythonD, open APIs, and pipeline processing, making it suitable for integrating internal business metrics, operations script results, and specialized technology stack data into Guance.