home / blog / yaml vs json

YAML vs JSON: What's the Difference?

They store the exact same kind of data, but one uses braces and the other uses whitespace you can't see, which is exactly why a misplaced indent in your Docker Compose file breaks the build without so much as a syntax error.

Ask someone to describe what JSON and YAML actually store, and you'll get the same answer for both: objects (or maps, if you prefer), arrays, and a handful of scalar types, meaning strings, numbers, booleans, and null. That's not a coincidence. YAML was designed from the start as a superset of JSON's data model, which is why any valid JSON document is also technically valid YAML. Rename a .json file to .yaml and a compliant parser will read it without complaint, indentation and all, because bracket-delimited JSON is just one of the many ways YAML allows you to write the same underlying tree of values.

So when people ask "yaml vs json, which is better," the honest answer is that the question usually points at the wrong layer of the problem. The data underneath doesn't change between the two. What changes is how much syntax you have to type to represent it, and how forgiving that syntax is when a human is the one typing it rather than a machine generating it automatically. JSON was built to be a clean, minimal wire format that machines could parse quickly and unambiguously, with no room for interpretation. YAML arrived a few years later specifically to fix JSON's biggest weakness for humans: reading and writing it by hand gets tedious fast once a document runs past a few dozen lines. Both design goals were achieved. Neither format is objectively wrong; they're built for different primary readers, and that distinction explains almost everything else about how they're used in practice.

This matters more than it sounds like it should, because the choice between the two isn't really a matter of taste. It's a question of who, or what, is going to be staring at the file most often: a parser running in a CI pipeline, or an engineer scrolling through it at 5pm trying to figure out why a deployment didn't roll out the way they expected.

Strict brackets vs. significant whitespace

JSON's syntax is small and rigid. Every object gets curly braces, every array gets square brackets, every key is a quoted string, every string value is quoted, commas separate everything, and a trailing comma is a parse error, full stop. There's exactly one way to write a given piece of data, no stylistic wiggle room. That rigidity is genuinely a feature, not a limitation: a JSON parser either accepts your document or rejects it with a specific line and column number, and two unrelated JSON serializers written in two different languages will produce output that's structurally identical aside from whitespace. You can validate JSON with a fairly small state machine, which is a big part of why it's implemented natively in essentially every language in use today and runs natively inside JavaScript itself.

YAML throws almost all of that punctuation away and replaces it with indentation instead. A key's children are whatever's indented further than the key itself, nothing more explicit than that. A list item starts with a dash and a space. Strings usually don't need quotes at all unless they contain something ambiguous. The result reads a lot closer to plain English than to code, which is exactly why config files gravitated toward it over the years. Nobody particularly enjoys hand-editing a two-hundred-line Kubernetes manifest full of nested brace pairs and trailing commas.

But that same convenience is also YAML's biggest footgun, and it's worth being blunt about why, because this is where most real-world pain comes from. In JSON, get the structure wrong and you get a parse error; the document simply refuses to load, and you know immediately something's broken. In YAML, get the indentation wrong and you usually still get a document back. It just parses into a different structure than the one you intended, silently. Indent a key two spaces instead of four and it quietly becomes a sibling of its old parent instead of a child. Nothing throws an error. Your YAML file is still technically "valid"; it's just valid for a shape you never meant to write in the first place. This single behavior is the root cause behind a huge share of "why is my deployment behaving strangely" tickets: nothing crashed, the YAML parsed fine on the surface, it just parsed into the wrong tree and nobody noticed until something downstream broke.

What YAML can do that JSON simply can't

Once you think of YAML as "JSON's data model plus a friendlier grammar," the extra features start to make sense as things a config-file format needs that a pure data-interchange format doesn't. Take comments: JSON has no comment syntax at all, which is a real limitation once you're maintaining a file six months after writing it and need to leave a note explaining why some value is set the way it is. YAML supports full-line comments with a leading #, and that one addition explains a lot of its adoption for anything a human is expected to edit and re-edit over time rather than write once and forget.

YAML also offers two ways to write multi-line strings without escaping every newline as \n. There's the literal block scalar, written with a pipe (|), which preserves line breaks exactly as written, and the folded block scalar, written with a right angle bracket (>), which joins lines into a single string separated by spaces instead. Anyone who's embedded a shell script inside a GitHub Actions workflow file has almost certainly used the pipe form without knowing its formal name; it's what lets you write ten lines of bash under a run: key and have it actually come out as ten readable lines rather than one unreadable blob of escaped text.

Then there are anchors and aliases, which behave more like a lightweight templating feature than a data-format feature. Define a block once with &base, reference it again elsewhere with *base, and the parser expands it inline at load time. That's genuinely useful in a Docker Compose file where several services share the same environment variables or the same logging configuration, and pasting that block five times would just invite drift between copies. On top of that, YAML allows multiple documents inside one file, separated by a bare --- line on its own, which is exactly how a Kubernetes manifest bundles a Deployment, a Service, and a ConfigMap into a single file that kubectl apply -f reads back as three separate objects. None of this has a real JSON equivalent, because JSON was never meant to be hand-authored at this scale in the first place. It was meant to be generated by code and parsed by code, with a human rarely in the loop.

Where each format actually wins, and the gotchas that catch people

In practice the split isn't really much of a debate; it mostly settles itself by context alone. JSON dominates anywhere a machine is both writing and reading the data with no human in between: REST and GraphQL API payloads, browser localStorage, npm's package.json, anything touched by JavaScript's native JSON.parse and JSON.stringify. It's fast to generate, fast to parse, and unambiguous between systems written in completely different languages, which is exactly what you want sitting at a network boundary.

YAML dominates anywhere a human is the primary author and a machine only reads it second: Kubernetes manifests, Docker Compose files, GitHub Actions and GitLab CI workflows, Ansible playbooks, and most application config files that aren't already tied to a specific language's own native format. The common thread running through all of those is that they're files people write by hand, review inside pull requests, and edit under time pressure, which is precisely the audience JSON's punctuation was never designed to serve well.

Two gotchas are worth knowing before you convert json to yaml and assume the result will just work everywhere without a second look. First, YAML's spec explicitly forbids tabs for indentation; spaces only, no exceptions. This trips up people whose editors default to inserting tabs, and the resulting failure mode is often a cryptic "mapping values are not allowed" error that doesn't mention tabs at all, which makes it harder to diagnose than it should be. Second, and the far more notorious one, is what the YAML community has nicknamed the Norway problem. Under the YAML 1.1 spec, which is what a surprising number of real-world parsers still implement by default (including PyYAML's default loader and older versions of Ruby's Psych), a small set of bareword strings get silently converted to booleans or other types if you leave them unquoted. The ISO country code for Norway, NO, happens to be one of them: write country: NO in a YAML file meant to hold country codes, and a YAML 1.1 parser reads it back as the boolean false rather than the two-letter string "NO", with zero warning that anything went wrong. The same thing happens with yes, on, and off, all of which YAML 1.1 treats as boolean literals unless you wrap them in quotes. This has shipped as a real bug in real data pipelines that ingest country-code lists, and it's exactly the kind of thing that only surfaces once someone in Norway's data goes mysteriously missing from a report. The fix is simple once you know it exists: quote any string that could plausibly be mistaken for a boolean, a number, or null, and don't lean on YAML's automatic type inference for anything that originated as user input.

Try it

GlaeKit's JSON to YAML / YAML to JSON converter handles both directions in your browser, including anchors, aliases, and multi-document files. Paste YAML alone and convert it to JSON and back to reformat and validate it in one pass. Nothing is uploaded anywhere.

Frequently asked questions

Can YAML represent anything JSON can?

Yes. YAML's data model is a superset of JSON's — objects, arrays, strings, numbers, booleans, and null all map directly. Any valid JSON document is also valid YAML as-is, which is why converters can round-trip between the two without losing information.

Why does my YAML file fail with an indentation error?

Almost always because a line is indented to a level that doesn't match any existing block, or because two sibling items use inconsistent indentation. YAML has no closing brackets to signal where a block ends, so the parser relies entirely on whitespace to figure out structure — a single space out of place is often the whole problem.

What is the "Norway problem"?

Under the YAML 1.1 spec that most real-world parsers implement, certain unquoted words get auto-converted to other types. The country code "NO" is read back as the boolean false, and words like yes, on, and off get parsed as booleans instead of strings. Quoting the value ("NO" instead of NO) avoids it entirely.

Can I use tabs in YAML?

No. The YAML spec disallows tab characters for indentation — only spaces are valid. A file with tab-indented lines will usually fail to parse, often with an error that doesn't mention tabs at all, so it's worth checking your editor's whitespace settings first if a YAML file suddenly won't load.

Should I use JSON or YAML for a new config file?

If a human will be editing it by hand, especially with comments explaining choices, YAML is usually the better fit, which is why most CI and orchestration tools default to it. If the file is generated and consumed entirely by code, with no human editing in between, JSON's stricter syntax and near-universal native parsing support make it the simpler choice.