JSONL vs JSON: Why Logs and LLM Datasets Use One Object Per Line

The difference is two characters and a comma, and it changes everything about how a file can be written, read and recovered.

JSONL vs JSON: Why Logs and LLM Datasets Use One Object Per Line
Share this article:

Open a training dataset or a structured log file and you will often find something that looks like broken JSON:

{"id": 1, "role": "user", "text": "Hello"}
{"id": 2, "role": "assistant", "text": "Hi there"}
{"id": 3, "role": "user", "text": "How are you?"}

No opening bracket, no commas, no closing bracket. Paste it into a JSON parser and it fails on line 2.

It is not broken. It is JSON Lines — JSONL, also called NDJSON. And the missing punctuation is the entire point.

What the difference buys you

Both formats hold the same records. The layout is what differs, and the layout determines what you can do with the file.

JSON arrayJSONL
Brackets`[ ... ]`None
SeparatorCommasNewlines
Append a recordRewrite the fileAppend one line
Read without full parseNoYes
One corrupt recordWhole file failsOnly that line fails
StreamingAwkwardNatural
Diff readabilityPoorGood

Appending is the big one. To add a record to a JSON array you must find the closing bracket, insert a comma, add the record, and rewrite the bracket. For a 50 GB file that means rewriting 50 GB. With JSONL you open the file in append mode and write one line.

Recovery is the second. If a process dies halfway through writing a JSON array, you have an unclosed bracket and an unparseable file. If it dies halfway through writing JSONL, you have a complete file plus one truncated last line — every earlier record is still readable.

Where you will meet it

  • LLM fine-tuning. OpenAI and most training APIs require JSONL for datasets. One example per line, appended as you generate them.
  • Structured logging. Each log entry is a line, so `tail -f` still works and log shippers can process entries independently.
  • Data warehouses. BigQuery, Snowflake and Athena all ingest newline-delimited JSON natively.
  • Container logs. Docker and Kubernetes emit JSONL when structured logging is on.
  • Streaming APIs. A long-running response can emit records as they are ready without knowing the total count in advance.

That last point is worth dwelling on: a JSON array requires you to know where the data ends. JSONL does not, which is why it suits anything unbounded.

Converting between them

You need a JSON array when a tool expects one — most JSON viewers, many libraries, and anything doing a whole-document transform. Use the JSONL to JSON converter in either direction.

The converter reports the line number of any record that fails to parse, which is the practical reason to use it rather than pasting into a general JSON tool. A generic parser can only tell you the file is invalid; with thousands of lines that is not actionable.

Working with JSONL on the command line

For files too large to paste, `jq` handles JSONL natively:

# Pretty-print every record
jq '.' data.jsonl

# Filter records
jq 'select(.role == "user")' data.jsonl

# Convert JSONL to a JSON array
jq -s '.' data.jsonl > data.json

# Convert a JSON array to JSONL
jq -c '.[]' data.json > data.jsonl

# Count records without parsing the whole file
wc -l data.jsonl

That last one is a nice illustration: counting records in JSONL is just counting lines. In a JSON array you have to parse the structure.

Common mistakes

  • Pretty-printing it. Formatting JSONL with indentation breaks it, because each record must occupy exactly one line. If your editor auto-formats on save, turn it off for `.jsonl` files.
  • Adding commas. There are none. A trailing comma makes each line invalid JSON on its own.
  • Wrapping it in brackets. That produces a JSON array, which JSONL readers reject.
  • Assuming uniform records. JSONL makes no such promise — records are independent, and heterogeneous records are normal in logs. If you need consistency, validate it explicitly.
  • Using `.json` as the extension. It misleads tools and people. Use `.jsonl` or `.ndjson`.

Is JSONL the same as NDJSON?

Effectively yes. NDJSON was a separate specification with slightly stricter rules about line endings and encoding, but in practice the terms are used interchangeably and tools accept both. If someone hands you a `.ndjson` file, treat it as JSONL.

Frequently asked questions

Can a JSONL record contain a newline?

Not a literal one. Inside a string it must be escaped as `\n`, which is what `JSON.stringify` does automatically.

Does JSONL need a trailing newline at the end?

Most tools accept either. Some strict readers want one, so including it is the safer default.

Is JSONL smaller than a JSON array?

Marginally — you save the brackets and commas. The advantage is operational, not about size.

How do I validate a large JSONL file?

`jq -e '.' data.jsonl > /dev/null` exits non-zero on the first invalid line. For a sample, the [JSONL converter](/jsonl-viewer) names the failing line.

Can records have different shapes?

Yes, and they often do. If you need to know what shapes exist, convert a sample to an array and run it through the [JSON Schema generator](/json-schema-generator).

Related reading

Once you have an array, the JSON Schema generator infers a schema across every record, which is a fast way to discover which fields are actually consistent. For TypeScript types from a single record, use the JSON to TypeScript converter.

Convert JSONL to JSON →