Skip to content
Datasets & storage

Dataset

What a mainframe calls a file: a named collection of records with a defined structure, rather than a loose stream of bytes.

Also written file, data set

A dataset is the mainframe equivalent of a file, and the difference in name reflects a real difference in design.

On most systems a file is a stream of bytes, and any structure inside it is the application's business. A dataset is a collection of records, and the system knows about that structure. When you create one you declare how long each record is, whether records are fixed or variable in length, and how much space to set aside. The system enforces it from then on.

That sounds like extra bureaucracy, and at first it is. The payoff is that utilities, sort programs and languages like COBOL all understand the layout without being told again, and the system can manage storage efficiently because it knows what is coming.

Datasets come in several shapes. Sequential ones hold records one after another. Partitioned ones act as a library of separately named members. VSAM datasets support keyed access for records retrieved by identifier rather than position.

Alongside all of these, z/OS also supports ordinary Unix files through Unix System Services.

Browse all 115 terms

Learn this properly.

Use Dataset for real in Mainframe101, in your browser, with Zed beside you. Join the waitlist.

Early access and updates. No spam, unsubscribe any time.