- Data collector
DataKit is an open-source, all-in-one data integration Agent that supports all operating systems (Linux/Windows/macOS), offering comprehensive data collection capabilities across hosts, containers, middleware, tracing, logging, and security inspections.
Collect hosts, containers, Kubernetes, middleware, databases, networks, logs, traces, and custom metrics Field extraction, type conversion, filtering, and desensitization are completed through Pipeline before inbound storage Compatible with ecosystem data sources such as OpenTelemetry, Prometheus, Telegraf, Fluentd, and StatsD
Data Collection Agent
What is DataKit?
DataKit is a Guance data collection Agent that can collect metrics, logs, traces, objects, events, and custom data from hosts, containers, Kubernetes, cloud hosts, and edge environments, and access data into Guance workspaces through pipelines, unified tags, and open protocols.
Functional features
Multi-dimensional data collection
Supports collecting data such as Metrics, Logs, and Traces from various infrastructures and technology stacks, and processing these data in a structured manner.

Kubernetes cloud-native technology support
Based on a cloud-native microservices architecture, it covers comprehensive data collection within the Kubernetes ecosystem.

Observable text processing pipeline
Built-in easy-to-use data extraction and processing engine, Pipeline, is used to extract unstructured data, making querying and statistics convenient.

Better tech stack support than Telegraf
Compared to Telegraf, which can only collect time-series data, DataKit covers a much broader range of data acquisition types, with simpler data collection configurations and better data quality.

Flexible deployment, simple and easy to use
DataKit is deployed and installed with one click, with hundreds of built-in data integrations, supporting automatic basic data collection during installation, ready to use out of the box.
Core Capabilities
From collection and processing to unified association
DataKit is not just a single collector. It handles data access, preprocessing, tag specification, ecosystem compatibility, and large-scale deployment and maintenance, bringing observable data into a unified analytical context.
Comprehensive data collection
Deploy DataKit across hosts, containers, Kubernetes, databases, middleware, and cloud resources, unifying collection metrics, logs, links, objects, and events, reducing the caliber differences caused by multiple collectors coexisting.
Pipeline data processing
Through Pipeline, log and text data are parsed, field extracted, typed, filtered, desensitized, and tag complete, making subsequent retrieval, alerting, and aggregation analysis more stable.
Open ecosystem compatibility
Supports data access from open ecosystems such as OpenTelemetry, Prometheus, Telegraf, Fluentd, and StatsD, helping existing collection systems smoothly enter Guance unified analysis.
Cloud-native deployment and governance
Supports deployment on Linux, Windows, macOS, Docker, Kubernetes, and cloud hosting, and can reduce large-scale maintenance costs through election, GitOps, DataKit API, and command-line tools.
Workflow
Recommended access methods
01Deploy DataKit first
Deploy DataKit on hosts, containers, Kubernetes, or cloud hosts, prioritizing access to infrastructure, critical applications, logs, and alert-related data.
02Standardize tags and pipelines
Unified service name, environment, version, region, team, and business tags, and use Pipeline to handle log fields, anonymization, and data quality.
03Enter Guance Unified Analysis View
Metrics, logs, links, RUM, events, and alerts are linked to the same object, forming a troubleshooting chain from problem phenomena to root cause location.
Ecosystem support 
DataFlux Func programmable data processing development platform

Supports receiving log data collected by Fluentd

Supports receiving metrics collected by Telegraf

Scheck cloud-native programmable security inspection plugin

Supports collecting and receiving Prometheus system metrics

Supports receiving Statsd metrics
FAQ
Frequently asked questions
What is DataKit?
DataKit is an Guance open-source, all-in-one data collection Agent used to collect data from hosts, containers, Kubernetes, logs, links, databases, middleware, networks, cloud resources, and custom business data.
Can DataKit be used together with OpenTelemetry, Prometheus, and Telegraf?
Yes, you can. DataKit can directly collect data or receive ecosystem data from OpenTelemetry, Prometheus, Telegraf, Fluentd, StatsD, and others, helping teams integrate existing collection links into Guance unified analysis view.
How does DataKit help logs and metrics perform correlation analysis?
During the collection phase, DataKit completes unified tags such as service name, environment, host, container, Pod, and version, and processes unstructured logs through pipelines, allowing metrics, logs, links, and alerts to be associated around the same object.
If you already have custom systems or internal metrics, can you access them via DataKit?
Yes, you can. DataKit supports custom data collection, PythonD, open APIs, and pipeline processing, making it suitable for integrating internal business metrics, operations script results, and specialized technology stack data into Guance.