The data category covers offensive work against the way information is represented and transformed, below the level of any single application. The same bytes flow through encoders, ciphers, and serializers on their way between systems, and each of those steps is an attack surface in its own right: a decoder that trusts its input, a cipher used without integrity, or a transform that is reversible when it was assumed not to be.
Why it is split this way#
The subcategories group by the kind of representation or transformation under attack, because the techniques and tools differ sharply between them:
- Encoding: reversible schemes that carry data without secrecy (Base64, URL, percent, hex). The offensive angle is smuggling, filter evasion, and parser confusion.
- Character encoding: charset and Unicode handling, where normalization, overlong forms, and homoglyphs defeat comparisons and validators.
- Ciphers: cryptographic transforms, attacked through misuse rather than math, for example missing integrity, predictable IVs, and oracle behavior.
- Binary data: raw and structured binary formats, hex-level manipulation, and format-aware tampering.
- Data transformation: compression, serialization, and format conversion, where the transform itself (not the payload) carries the flaw.
Splitting on representation keeps each page focused on one class of parser or algorithm, so a technique that abuses Unicode normalization does not get tangled with one that abuses a block cipher mode. These are foundational primitives that recur across the software, network, and application layers, which is why they live in their own category rather than under any one target.