Sanskrit adjusts its sounds at every junction, inside a word and between words, according to a set of rules the grammarians set out exhaustively. Vowels merge, consonants assimilate to what follows, a final consonant changes according to the first sound of the next word, and nasals and sibilants shift between their dental and retroflex forms under stated conditions.
The script records the result and not the input. What is written is the adjusted form, so a reader is always working backwards, and separating a line of Siddham into its words is an act of grammatical analysis rather than of recognition. In a script that also runs its words together without spaces, as Indian manuscript practice often does, the analysis is the reading.
That has a direct consequence for how the script could be taught in East Asia. A student who learned the letters alone could copy a Sanskrit text accurately, produce the syllables of a dharani correctly and recognise a seed syllable on an image, all of which were the point. Construing a sentence was another matter and required the grammar, which is why the Japanese tradition of Siddham study produced letter manuals in quantity and grammatical study only alongside them.
The distinction between copying a script and reading a language is real, and the shittan tradition sits deliberately on one side of it.