SRE Foundation Certification: Career Value, Cost and Study Plan
Updated on 10 October 2026. What’s new: exam fee corrected to $257 and the vendor's page lists the exam as open book, not closed book
If you carry a pager for servers someone else designed, a credential will not make you an SRE, but it can show a hiring manager that you already speak the language: service level objectives, error budgets, toil, blameless reviews. SRE Foundation from DevOps Institute is the entry paper for that language. It has 40 questions, runs 60 minutes, needs 65% and costs $257, and this guide explains what it is worth to your career and how to prepare for it.

DevOps Institute, which created the exam, is now part of the PeopleCert group. The vendor's exam page says so in a footnote, and its purchase button hands you over to PeopleCert's platform to pay. I read both the DevOps Institute page and PeopleCert's listing in October 2026. The exam is still on sale, and neither page shows a version number, so there is no newer edition to chase.
What does this credential tell a hiring manager about an ops engineer?
It tells them three modest things. First, that you have studied the vocabulary of reliability work in a structured way rather than picking it up from incident chatter. Second, that you can separate ideas that sound alike, such as an indicator, an objective and an agreement. Third, that you chose to spend evenings on this instead of waiting for your employer to send you on a course.
It does not tell them that you can design a monitoring stack, run a game day or negotiate an error budget with a product owner. Those skills are shown in interviews through stories from your own on-call weeks. The most useful way to treat the certificate is as a prompt for those stories: for every module you study, find one moment from your last year of operations work that illustrates it, and rehearse it in two minutes.
Who gets the most from it?
The vendor describes the audience broadly: developers, system administrators, IT managers and DevOps engineers all fit. In practice three groups gain most. System administrators and support engineers who want a route toward platform or reliability roles. Developers whose teams have just been handed on-call duty. And team leads who need a shared vocabulary before they start setting reliability targets with their stakeholders. Nothing in the vendor's page lists a formal prerequisite, so you can sit the paper without holding another certificate first.
Is there a Google SRE certification?
Many searches on this topic ask for a Google credential, so it is worth being clear. Site Reliability Engineering as a discipline was popularised by Google, and the books Google published about it are widely read. But SRE Foundation is awarded by DevOps Institute, not by Google. The two are related by ideas, not by ownership. If a job advert asks for a Google cloud credential, that is a different exam from a different vendor.
What are the exam facts today, and who runs the credential?
The table gives the figures you need before you plan anything. The fee is quoted in US dollars. The DevOps Institute page does not print a price at all, and PeopleCert's listing, which shows prices in pounds, displayed a starting price of £213.00 when I read it. What you actually pay depends on your currency, your country and whether you buy a training bundle, so confirm the amount at checkout.
| Item | Detail |
|---|---|
| Exam name | DevOps Institute Site Reliability Engineering Foundation |
| Exam code | SRE Foundation |
| Vendor | DevOps Institute, a member of the PeopleCert group |
| Questions | 40 |
| Duration | 60 minutes |
| Passing score | 65% |
| Exam fee | $257 |
| Open book | Yes, according to the vendor's page |
| Delivery | Web-based |
| Languages | English, Brazilian Portuguese, Chinese, French, Japanese, Spanish |
| Certification validity | 3 years, per the vendor |
Two details differ from what you may have read in older guides: the exam is open book, and the fee is the figure above. The question count, time limit and pass mark match on both vendor pages I checked, the DevOps Institute exam page and PeopleCert's SRE Foundation listing. Sixty-five percent of 40 questions is 26 correct answers, so you can miss 14 and still pass.
What does 90 seconds a question mean for an open-book paper?
It means the book is a safety net, not a strategy. Looking up a definition costs you the better part of a minute, and you only get 90 seconds on average. Candidates who plan to look everything up tend to finish with a handful of questions unanswered. Treat open book as permission to check one or two terms you half-remember, and prepare as if the paper were closed.
How do the vendor's modules map onto work you already do?
The vendor's exam page lists eight modules under what you will learn, and the vendor's blueprint graphic names nine themes. Neither document publishes a percentage per topic, so you cannot weight your revision by marks the way you can on some exams. The safest plan is to give every module a fair share of time and to spend extra on the ones that are new to you. The table below pairs each module with a place you have probably met it already, which is a quick way to find the gaps.
| Module on the vendor's page | Where an ops engineer has probably met it |
|---|---|
| SRE Principles and Practices | Arguments about whether a team should keep adding features or fix reliability first |
| Service Level Objectives and Error Budgets | An uptime promise in a customer contract that nobody can trace to a measurement |
| Reducing Toil | The weekly manual restart, ticket or certificate renewal that nobody has automated |
| Monitoring and Service Level Indicators | Dashboards that show CPU but not whether a user's request worked |
| SRE Tools and Automation | Deployment scripts, runbooks and pipelines that one person understands |
| Anti-Fragility and Learning from Failure | The outage that changed how the team works afterwards |
| Organizational Impact of SRE | Who owns on-call, and how reliability work gets onto a roadmap |
| SRE, Other Frameworks, The Future | How the work sits beside DevOps, ITIL and platform engineering in your company |
The blueprint graphic uses slightly different labels: culture, toil reduction, measurements, anti-fragility, SLAs/SLOs/SLIs, work sharing, deployments, performance management and incident management. Both lists describe the same territory, so learn the topics rather than the exact labels. The blueprint also states a few working figures, including a 50% ops and development load and a 25% on-call load, which come up again below. The syllabus overview on this site lists the same areas in one place if you want a checklist to tick off.
What is work sharing, and why does a percentage appear?
Work sharing is about stopping operations from eating the whole week. The blueprint mentions a 50% split between operational work and development work, and a 25% on-call load. The point is not the exact arithmetic but the principle: when toil and interrupts take more than their share, engineering time for fixing causes disappears and the team falls further behind. Expect a question that describes a team in that trap and asks which action restores the balance.
Where do indicators, objectives and error budgets trip people up?
These three ideas, together with the agreement that sits above them, are a reliable source of wrong answers because the words overlap. A service level indicator is a measurement, an objective is a target for that measurement over a period, and an agreement is a business contract with consequences if the target is missed. The error budget is whatever is left once you subtract the objective from 100%.
A worked example with real arithmetic
Suppose a checkout service defines its indicator as the share of requests that return a correct response within half a second. The team sets an objective of 99.9% over a rolling 28 days. A 28-day window has 40,320 minutes, so the error budget is 0.1% of that, or 40.32 minutes of bad service. If a bad deployment causes 25 minutes of failures on day nine, 15.32 minutes remain for the next 19 days. A sensible team now slows risky releases, because the remaining budget is thin; a team with most of its budget unspent can ship faster.
Now try the questions an exam writer might ask from that scenario. Which number is the indicator? The measured percentage of good requests. Which is the objective? The 99.9% target. What would turn it into an agreement? A customer contract that promises, for example, a credit when the objective is missed. Who decides how the remaining budget is spent? The product and engineering sides together, which is why the budget is a shared tool and not an operations weapon.
Four slips to check in your own head
- Treating a 100% objective as a sensible goal. It leaves no budget for change, and the cost rises steeply for every extra nine.
- Measuring what is easy rather than what users feel. CPU load is a signal about a machine; a failed checkout is a signal about a person.
- Confusing a window with a deadline. The period over which the objective is judged is part of the definition.
- Using the agreement as the target. Agreements are usually set looser than the internal objective so the team has a margin before penalties begin.
What does a calm incident look like from first alert to review?
The incident management area rewards candidates who can describe a structured, unheroic response. The lifecycle has five stages: detection, response, remediation, analysis and readiness. The stages run in a loop, because readiness work, such as better alerts and rehearsed runbooks, is what makes the next detection faster. The graphic shows the path as a roadmap.

Who does what during the response?
Three roles come up repeatedly. The incident commander coordinates and decides, but does not type the fix. The communications lead keeps stakeholders informed so engineers are not interrupted for status. The operations lead works on the system itself. Separating the roles is the whole idea: it stops the loudest or most senior person from doing everything at once.
Why is the review blameless?
Because blame stops people reporting what really happened. A blameless review asks why the system allowed a mistake to cause harm, so the fixes are about guard rails, alerts and automation, not about telling one person to be more careful. In an exam, an answer that names a person as the root cause is almost always the wrong one; the right one describes a process or tooling gap and a follow-up action with an owner.
Where do anti-fragility and chaos engineering fit?
Robust systems resist failure. Anti-fragile ones are built on the assumption that failure will come, and they improve under stress. Chaos engineering, the habit of injecting faults such as a terminated server or added latency on purpose, is the practical form of that idea. The related term blast radius means how much of the service a single failure can damage, and gradual releases such as canary and blue-green deployments exist to keep it small.
Does open book let you skip learning the material?
No, and the clearest reason is the style of question. Many items describe a situation and ask which practice fits, so a keyword search of a book does not hand you the answer. Practise on classification, because it is the skill the exam keeps returning to. Take this list of recurring tasks in an operations team and decide which count as toil, meaning manual, repetitive, automatable work that grows with the service and leaves no lasting value.
- Restarting a stuck worker by hand each Monday. Toil: manual, repetitive and a script could do it.
- Writing a script that restarts the worker automatically. Not toil: it is engineering work that removes toil.
- Granting the same access request by hand forty times a month. Toil: it scales with demand and can be self-service.
- Investigating a new, unexplained latency spike. Not toil: it needs judgement and teaches the team something.
- Copying release notes into a ticket after every deploy. Toil: repetitive and can be generated from the pipeline.
If you cannot sort a list like that quickly, more reading will not fix it; practice with explained answers will. The same applies to deployment choices. Know when a canary release limits risk, when blue-green gives a fast rollback, and when a feature flag lets you separate deployment from release.
How can a working engineer fit preparation around on-call weeks?
A fixed calendar fails the moment an incident lands on your week, so build slack into the plan. The four-week timeline below assumes about five hours a week and keeps the final days as a buffer. If you are on call in week two, move that week's reading into the other weeks rather than skipping it.
- Week 1, principles and culture. Read the module list on the vendor's page, then write one sentence per module in your own words. Work on the indicator, objective and agreement distinction until you can explain it to a colleague.
- Week 2, measurement and budgets. Redo the 28-day error budget calculation with three different objectives. Look at one of your own dashboards and ask which panels are true indicators.
- Week 3, toil, automation and incidents. Audit your last month of tickets for toil. Read one public postmortem and label each stage of the lifecycle.
- Week 4, timed practice and review. Sit a full 40-question paper in 60 minutes, list every miss by module, and revisit only those modules. A set of explained questions such as an SRE Foundation practice exam helps here, because the reasoning behind each answer is what closes a gap.
Book the paper once your timed scores sit comfortably above 65% for two attempts in a row. A single lucky result is not a reason to book. Because the exam is web-based, you can choose a quiet slot that does not collide with a rota.
Which study sources help, and which ones put your certificate at risk?
Start with the vendor. The exam page links a course description and the blueprint, and DevOps Institute lists instructor-led training, online learning and self-study as routes. Then add the public reading that SRE grew out of: the books and articles Google has published on the subject, the postmortems that large companies share after major outages, and the engineering blogs of firms that run services at scale. When you read, keep asking which module a passage illustrates and what the indicator or objective was.
Why avoid exam dumps?
Searches for dumps and leaked questions are common for this exam, and the answers they carry are often wrong or out of date. Memorising them teaches you the wrong things about topics such as error budgets, and it breaks the certification body's rules, so you risk a revoked result. A genuine pass is worth something in an interview only if you can discuss it, and a dump leaves you with nothing to say.
Is a study group worth the time?
Yes, if it is small and regular. Pair with a colleague, take turns explaining a module in five minutes, and quiz each other on the confusing pairs: toil and engineering work, objective and agreement, robust and anti-fragile. Teaching a concept exposes the gaps that reading hides. Online communities and meetups can help, but treat their advice as colour, not as exam fact, and check anything about the format against the vendor's own pages.
What should you do in the week after you pass?
Put the credential where recruiters look, then turn it into evidence. Pick one piece of toil in your current job, measure it for a fortnight, automate or remove it, and write down the before and after. That single story, told with an indicator and an objective, will carry more weight in an interview than the certificate alone, and it gives your next career conversation something concrete to build on. The certificate is valid for three years, according to the vendor, so there is time to add to it.
Frequently Asked Questions
Is there a Google SRE certification?
SRE Foundation is not a Google credential. It is awarded by DevOps Institute, which is part of the PeopleCert group. Google popularised the discipline through its published books, but the exam, its 40 questions and its 65% pass mark all come from DevOps Institute.
How much does the SRE Foundation certification cost?
The exam fee is $257. The DevOps Institute page shows no price and sends buyers to PeopleCert's platform, whose listing showed a starting price in pounds in October 2026. The amount you pay depends on your currency and any training bundle, so confirm it at checkout.
How many questions are on the SRE Foundation exam, and what score passes?
The exam has 40 questions in 60 minutes, and 65% is needed to pass. That is 26 correct answers out of 40, so 14 misses are allowed. The vendor delivers it on a web-based platform in six languages, including English, Japanese and Spanish.
Is the SRE Foundation exam open book?
Yes, the vendor's exam page lists the exam as open book. With 40 questions in 60 minutes you have about 90 seconds each, so searching for every answer is not realistic. Learn the material and use the book only to check a term you half-remember.
How long is the SRE Foundation certification valid?
The DevOps Institute exam page states that the certification is valid for three years. DevOps Institute is a member of the PeopleCert group and purchases are completed on PeopleCert's platform, so check the vendor's current renewal terms before the three years end.
- DevOps Institute Site Reliability Engineering Foundation Test Questions |
- SRE Foundation Question Bank |
- Site Reliability Engineering Foundation |
- Site Reliability Engineering (SRE) |
- Site Reliability Engineering Foundation Simulator |
- Site Reliability Engineering Foundation Mock Exam |
- DevOps Institute Site Reliability Engineering Foundation Book |
- Site Reliability Engineering Foundation Certification Cost |
- Site Reliability Engineering Foundation Certification Requirements |
- DevOps Institute Site Reliability Engineering Foundation Sample Questions |
- SRE Foundation Exam Questions Download |
- SRE Foundation Test Questions |
- Site Reliability Engineering Foundation PDF
