HL7 Interoperability

Anatomy of an HL7 v2 Message

  • 5 min
  • 7 steps
  • 2 questions
  • Lesson 2 of 51

In this lesson

  1. A message is an ordered tree
  2. MSH is deliberately special
  3. Locate data precisely
  4. Message type, event, and structure are different
  5. Segment grammar matters
  6. Empty, null, and omitted are distinct
  7. A safe parsing sequence
HL7 Message StructureHL7 SegmentHL7 Encoding CharactersTrigger Event

A message is an ordered tree

HL7 v2 looks like plain text, but its structure is hierarchical:

message
└── segment
    └── field (may repeat)
        └── component
            └── subcomponent

Segments are normally terminated by carriage return. Within each segment, delimiters separate the lower levels. The standard’s Control chapter defines these generic construction rules for all v2 messages 1.

Here is a compact admission message:

MSH|^~\&|ADT1|HOSP|LAB|HOSP|202406011230-0500||ADT^A01^ADT_A01|MSG00001|P|2.5
EVN|A01|202406011230-0500
PID|1||100711^^^HOSP^MR||DOE^JANE^Q||19800101|F
PV1|1|I|2000^2012^01||||004777^SMITH^JOHN

Read it top to bottom: message header (MSH), event (EVN), patient identity (PID), and visit (PV1). The order and repetition of segments are defined by the message structure—not chosen arbitrarily 2.

The anatomy of an ADT^A01 admit message: segments form an ordered structure, and each segment contains fields interpreted using the delimiters declared in MSH.
The anatomy of an ADT^A01 admit message: segments form an ordered structure, and each segment contains fields interpreted using the delimiters declared in MSH. source

MSH is deliberately special

Most segments begin with a three-character name followed by the field separator. MSH must also declare the separators needed to parse the rest of the message:

MSH|^~\&|...
   │││││
   ││││└─ subcomponent &
   │││└── escape \
   ││└─── repetition ~
   │└──── component ^
   └───── field |

The field separator is MSH-1; the next four characters are MSH-2. Do not split the full MSH line using assumed defaults before reading these positions. A robust parser discovers the delimiters from the message itself and then applies them consistently.

Locate data precisely

Interface teams commonly identify an element as segment-field.component.subcomponent, optionally naming a repetition. Examples:

  • MSH-9.1 = message code, here ADT;
  • MSH-9.2 = trigger event, here A01;
  • MSH-9.3 = message structure, here ADT_A01;
  • PID-3[1].1 = first repetition’s identifier value;
  • PID-3[1].4 = its assigning authority;
  • PV1-3.1 = point of care within the patient location.

This notation is more useful than saying “the patient ID field.” PID-3 can repeat, and each CX repetition contains its own assigning authority and identifier type. A value without its structural context is often ambiguous.

The data-type chapter defines the components that make a composite meaningful 3. For example, DOE^JANE^Q is not three unrelated strings; it is an XPN name whose positions carry family, given, and second-name roles.

Message type, event, and structure are different

MSH-9 may contain three components:

ADT^A01^ADT_A01
  • ADT names the broad message code;
  • A01 names the real-world trigger event;
  • ADT_A01 names the abstract segment structure.

Several events can share one structure. Routing only on ADT throws away important workflow meaning; routing on the full declared contract is safer. Likewise, a receiver should not infer event meaning only from which optional segments happened to appear.

Segment grammar matters

A message definition behaves like a grammar: some segments are required, some optional, and some repeat in groups. An order message may repeat an ORC/OBR group; a result may repeat observations beneath an order. Flattening the message into an unordered dictionary loses these relationships.

Three visual cues are not enough by themselves:

  • a segment appearing once does not prove its maximum cardinality is one;
  • an absent optional segment is not the same as a prohibited segment;
  • a parseable segment in the wrong group can still be nonconformant.

Validate the actual message against the agreed version and profile, not merely against a generic list of segment names.

Empty, null, and omitted are distinct

Two separators with nothing between them indicate an unpopulated field. In many v2 contexts, a pair of double quotes ("") is used as a null value instruction. Omission, empty content, and explicit null can drive different update behavior under an implementation guide. Never normalize them all to the same database value without a documented rule.

The same caution applies to trailing fields and components. Senders often omit unused trailing delimiters. A positional parser must preserve the intended field numbers even when the textual line looks short.

A safe parsing sequence

  1. Preserve the original bytes and message boundaries.
  2. Confirm the first segment and discover MSH delimiters.
  3. Split segments using the segment terminator.
  4. Parse fields, repetitions, components, and subcomponents without discarding empty positions.
  5. Read version and message type from MSH.
  6. Select the agreed message profile.
  7. Validate structure, usage, cardinality, length, data types, and value sets.
  8. Map into application data only after validation and identity checks.

Logging should identify the message and error location without casually duplicating protected health information. Preserve a secure audit trail that connects the inbound control ID, acknowledgment, route, transformation version, and downstream outcome.

Worked reading exercise

For the sample message, answer:

  1. What event occurred, and which structure was declared?
  2. Which organization assigned identifier 100711?
  3. Is 2000^2012^01 one field or three fields?
  4. Which time values include an explicit UTC offset?
  5. Which assumptions would still require an interface profile?

Then deliberately mutate one element: remove the assigning authority, place A03 in the event component without changing the workflow, or add an unexpected PID repetition. Predict whether the failure should be caught by parsing, profile validation, semantic validation, or application reconciliation. That separation is the foundation of useful interface troubleshooting.

Practice

What determines the field separator in a v2 message?

Practice

What is the safest interpretation of two consecutive field separators?

Lesson complete

Nice work.

1day streak
0/1today's goal
–correct

Up next · 5 min

How v2 Messages Are Exchanged — Acknowledgments and MLLP

Next lesson
Sources for this lesson
  1. 1
    HL7 Version 2.9 — Chapter 2: Control. HL7 International (HL7 Europe public mirror). 2019. verifiedDefines v2 message construction, delimiters, message control, original and enhanced acknowledgment modes, MSH, MSA, ERR, and processing rules. Cited at: message construction and control.
  2. 2
    HL7 Standards — Section 1d: Version 2 (V2). HL7 International. verifiedThe HL7 Version 2 messaging standard, first released October 1987 and the most widely implemented healthcare messaging standard worldwide. Cited at: message definitions.
  3. 3
    HL7 Version 2.9 — Chapter 2A: Data Types. HL7 International (HL7 Europe public mirror). 2019. verifiedDefines primitive and composite v2 data types, including CWE, CX, XCN, XPN, identifiers, names, timestamps, coded values, and their components. Cited at: composite data types.

Further reading