Gremlin provides enterprise reliability management software for engineering organizations that need to uncover reliability risks before they become outages. Its platform is built to help teams standardize resilience practices, replace backward-looking incident metrics with forward-looking reliability measurement, and validate how services behave under real failure conditions across infrastructure and application layers.
Gremlin supports cloud, hybrid, and on-prem environments, including Kubernetes, containers, virtual machines, Windows, Linux, and serverless workloads. The company’s offering is aimed at organizations that need a repeatable way to test redundancy, scalability, dependency behavior, and disaster recovery readiness while giving engineering leaders a clearer view of reliability posture across teams and services.
Offerings, Capabilities, and Integrations
Gremlin combines proactive resilience testing, automated reliability measurement, risk detection, dependency mapping, and failover validation in a single platform. It enables teams to run scheduled and event-driven tests, monitor service health during tests, prioritize remediation work, and track reliability progress over time without relying solely on post-incident analysis.
Its capabilities span infrastructure-level and application-level testing, with deployment options for SaaS and isolated private-network environments. Gremlin also supports integrations with cloud platforms, observability systems, collaboration tools, and workflow systems so teams can connect testing with monitoring, alerting, ticketing, and operational review processes.
Products and Services
- Reliability Management: Gremlin’s core platform for standardizing reliability programs with prebuilt and customizable test suites, scheduling, scoring, and service-level visibility across engineering teams.
- Chaos Engineering: A controlled chaos engineering offering for safely recreating outages and validating how systems respond to infrastructure, network, and application failures.
- Private Edition: A fully isolated deployment of Gremlin that runs inside a customer-controlled private network, including cloud-hosted, datacenter, and air-gapped environments.
- Fault Injection: A fault injection capability with experiments for compute, network, and state failures across services, hosts, and containers.
- Reliability Scoring: A scoring capability that converts test outcomes into an objective service-level reliability score and trend view for ongoing measurement.
- Detected Risks: An automated risk detection feature that surfaces high-priority reliability issues such as misconfigurations, anti-patterns, and missing safeguards without requiring active tests.
- Dependency Discovery: A dependency mapping feature that automatically discovers and tracks service dependencies so teams can test and manage upstream and downstream reliability risks.
- Failure Flags: An application-level testing capability for Kubernetes, serverless, and managed environments that enables fault injection without requiring broad infrastructure access.
- Reliability Intelligence: An analysis and recommendation layer that interprets failed tests, identifies likely root causes, and provides guided remediation steps, with optional MCP-based AI workflows.
- Disaster Recovery Testing: A disaster recovery validation capability for safely testing zone, region, and datacenter failover scenarios across multiple services and teams.
- Scenarios: Multi-step reliability workflows that combine experiments and health checks to simulate real incident patterns and document test outcomes.
- Intelligent Health Checks: Automatically generated health checks that monitor traffic, error, and latency signals during testing for supported cloud and application environments.
Target Customers
Gremlin primarily targets enterprise engineering organizations running business-critical digital services. Its typical buyers and users include site reliability engineering, platform engineering, DevOps, performance engineering, and infrastructure teams that are responsible for uptime, release confidence, migration safety, and operational resilience.
The platform is especially relevant for companies operating distributed systems across AWS, Microsoft Azure, and Google Cloud, as well as Kubernetes, containers, and serverless environments. It is well suited to organizations in sectors such as financial services, retail and ecommerce, SaaS, and other high-availability environments where downtime, failed launches, and untested disaster recovery plans carry meaningful business risk.
Cloud Integrations and Marketplace
- AWS Marketplace: Gremlin is available through AWS Marketplace and supports AWS account integration for Elastic Load Balancer service mapping, Auto Scaling Group preparation, CloudWatch-based health checks, and Intelligent Health Checks.
- Azure Marketplace: Gremlin is available through Azure Marketplace and supports Azure Intelligent Health Checks for services such as Application Gateway, Front Door, App Service, Container Apps, and Application Insights.
- Google Cloud Platform: Gremlin supports Google Cloud Platform through Intelligent Health Checks that use Google Cloud Monitoring for services such as Cloud Run and components in the Google Cloud load balancing stack.
Key People
- Kolton Andrus: CEO
- Aaron Kaffen: VP of Marketing
- Lorne Kligerman: Director of Product
- Tammy Butow: Principal SRE
- Jason Yee: Director of Advocacy
- Ryan Detwiller: Director of Product Marketing
- Andre Newman: Sr. Reliability Specialist
Key Facts
- Headquarters: San Jose, California, United States
- Employees: Approximately 70
- Annual Revenue: Undisclosed
- Parent Company: None
- Subsidiaries: None
- Publicly Listed: No (privately held)