Skip to main content

CUSTOMER STORY

Reliable systems for
scientific discovery.

ABOUT XTALPI

About XtalPi

Founded in 2015, XtalPi combines quantum physics, artificial intelligence and robotic automation to support drug and materials research. Its work connects computational research with laboratory automation, creating an ongoing need for reliable computing jobs, clusters and application services. This case focuses on the technical platform and operations supporting that work.

AI research Industry
2015 Founded
Shenzhen One of its key operating bases, China

BUSINESS CHALLENGES

Growing workloads, fragmented investigation context

Computing jobs, cluster components and application services all support research workflows. When a job fails or a service slows down, engineers need to distinguish resource pressure from task failures and application issues.

Separate tools held logs, traces and metrics without continuous context. Investigations often stalled at the handoff between systems.

Multiple clusters and compute pools needed a shared health view. Aggregate capacity alone could not explain job status, GPU utilization or their relationship to application services.

Detecting an issue was only the beginning. Teams also needed severity-based alerts, follow-up and availability checks to coordinate investigation and recovery.

KEY RESULTS

Several×

Faster fault localization and root-cause analysis

Seconds

Business and compute-pool metric monitoring

Unified

Multi-cluster and application observability

SOLUTION

XtalPi × Guance — Solution

01

See clusters and compute pools in one place

Bring cluster health, compute-pool jobs and GPU utilization together with business metrics. Engineers can move from the overall operating picture to an individual job, using the same evidence for capacity discussions and incident investigation.

3

Focus areas: cluster health, compute jobs, GPU utilization

02

Unify logs, traces and metrics

Integrate existing telemetry and progressively consolidate fragmented data into one platform. Preserving object, time and call relationships reduces repeated searches, helping make fault localization and root-cause analysis several times faster.

2

Faster investigation workflows: fault location and root-cause analysis

03

Connect detection, collaboration and verification

Severity-based alerts, inspections and incident tracking support follow-through, while synthetic tests check availability from the access side. Development and operations share evidence from the first alert through investigation and verification.

4

Connected steps: monitor, alert, investigate, verify

BUSINESS IMPACT

Let platform teams focus on supporting research

01

Get from an alert to useful evidence sooner

Unified search and correlation reduce tool switching and support faster localization and root-cause investigation.

02

Ground resource decisions in actual work

Job status and GPU utilization add task-level context to compute-pool health and capacity discussions.

03

Keep checking application availability

Cluster monitoring and synthetic tests provide complementary evidence for alerting, investigation and verification.

Reliable foundations for scientific progress

Bring compute pools, GPUs, jobs and application traces into an observability model suited to research platforms.