All human-first jobs

Incident Operations Specialist

🏢 Zapier📍 NAMER📅 1 month ago💼 Full-time🌐 Remote

👥 Human review score

70%

Likelihood a person reads your application

📞 Interview rate

55%

Applicants who reach a human interview

Free · no account needed

Will your résumé get past the filter for this job?

Get an instant match score against this exact posting, plus the keywords you are missing.

Check my résumé — free

About the role

AI AT ZAPIER At Zapier https://zapier.com/about, we build and use automation every day to make work more efficient, creative, and human. So if you’re using AI tools while applying here - that’s great! We just ask that you use them responsibly and transparently. Check out our guidance on How to Collaborate with AI During Zapier’s Hiring Process https://zapier.com/l/jobs/ai-at-zapier, including how to use AI tools like ChatGPT, Claude, Gemini, or others during our hiring process - and when not to.   Job Posted: July 10th, 2026 Location: NAMER HI THERE! As Zapier expands into the enterprise market and accelerates AI-driven development, incident management is increasingly critical to customer trust and operational reliability. The Incident Operations Specialist keeps that program running day to day through reliable tooling, clean data, repeatable workflows, and AI-powered automation. You’ll report to the Incident Program Manager and help shape how Zapier responds to incidents, learns from them, and supports the people doing that work. This is an operations role with technical depth, not a software engineering role. We care most about two things: proven incident management experience and genuine AI fluency. The rest is coachable. - Our Commitment to Applicants https://zapier.com/jobs/our-commitment-to-applicants/ - Culture and Values at Zapier https://zapier.com/jobs/culture-and-values-at-zapier/ - Zapier Guide to Remote Work https://zapier.com/learn/remote-work/ - Zapier Code of Conduct https://zapier.com/jobs/zapier-code-of-conduct/ - Diversity and Inclusivity at Zapier https://zapier.com/jobs/working-on-diversity-and-inclusivity/ WHAT WE'RE LOOKING FOR AI fluency (required, not optional). This is a hard requirement at point of hire, not something you'll grow into on the job. Concretely, we're looking for: - You use AI-native tools (Cursor, Claude, Copilot, or similar) as your default working environment, not as a novelty. - You've built AI-powered workflows that keep running when you're offline, not one-off prompts. You can describe two or three specific examples, what they replaced, and what verification you built in. - You can quantify how AI has changed your throughput or quality. - You know when AI output needs checking, especially under incident-time pressure, and you have a point of view on how you calibrate trust. - Please note: If your AI usage is mostly occasional prompting of a chat interface, this role isn't the right fit yet. Incident response and analysis experience: You've worked in incident response, reliability, or a closely adjacent role. You've been hands-on with tools like incident.io http://incident.io and PagerDuty: on-call rotations, escalation paths, routing, integrations. You've analysed incidents after the fact, spotted patterns across many of them, and turned that into program-level improvements. Technical depth to build your own tools: You're not a software engineer, but you can write SQL against Databricks, wire up API integrations, build Slack workflows, and prototype lightweight AI agents. If a workflow doesn't exist, you build it. If a dashboard is broken, you fix it. How you work: Async-first and visible: status in public channels, no need to chase. You close the loop, prioritise ruthlessly, and push back on off-program requests rather than getting pulled thin. You translate technical detail into plain language for Support, GTM, and leadership without losing the signal. WHAT YOU'LL DO - Own incident tooling operations. Keep incident.io http://incident.io, PagerDuty, on-call rotations, escalation paths, and Slack-based workflows configured, reliable, and integrated. Fix what breaks. - Build and maintain AI-powered workflows. Thread summarisation, postmortem drafting, follow-up triage, severity classification, data hygiene. Turn one-off experiments into durable systems. - Analyze incidents and drive improvement. Participate in incidents and postmortems, spot recurring patterns, surface program-level friction with recommended fixes, not just problems. - Operate data and reporting. Build and troubleshoot dashboards and reports (Databricks, Grafana, Looker). Guard data quality and metric accuracy. - Sustain the IC community. Grow the community of practice for Incident Commanders and Support Leads. Coach responders on what good looks like. - Keep documentation usable under pressure. Playbooks, templates, guides. Flag gaps where program-level guidance needs updating. OUR STACK - Incident: incident.io http://incident.io, PagerDuty, Slack - Data and observability: Databricks, Grafana, Looker, SQL, Datadog, Prometheus, Opensearch, Graylog - AI: Cursor, Zapier AI, Claude, or equivalent - Collaboration: GitLab, Coda, Google Workspace, Jira, Zendesk APPLICATION DEADLINE: The anticipated application window is 30 days from the date job is posted, unless the number of applicants requires it to close sooner or later, or if the position is filled. Even though we’re an all-remote company, we still need to be thoughtful about where we have Zapiens working. Check out this resource https://docs.google.com/spreadsheets/d/1_lXQvKwuU4xXEbye7EW_En8iv1ZIMlQNmT7B2Mzbo24/edit?gid=0#gid=0 for a list of countries where we currently cannot have Zapiens permanently working.

Similar human-first roles