A Japanese monk reading a Siddham syllable does not pronounce it as a Sanskrit speaker would, and the tradition has never pretended otherwise. Japanese has no retroflex series, no voiced aspirates, no distinction between the three sibilants and no vocalic liquids, so a large part of the inventory the table sets out cannot be produced by a speaker trained only on Japanese sounds.
What developed instead is a set of reading conventions: an established Japanese pronunciation for each syllable, transmitted with the letters, which maps the Sanskrit inventory onto sounds a Japanese speaker can make. Distinctions that cannot be carried are merged, and the merger is systematic rather than accidental, so a syllable has a settled reading that a student learns along with its shape.
This is not a failure of the transmission and is not treated as one. The letters preserve the distinctions exactly, on the page, whether or not any given reader can voice them, which means the written record remains recoverable even where the spoken one has flattened. A scholar working from a Japanese Siddham manuscript reads the Sanskrit inventory from the letters and the Japanese reading tradition from the annotations, and the two are separate pieces of evidence about separate things.
The same arrangement holds for Sanskrit read anywhere outside India, and for Latin read anywhere in medieval Europe.