How this implementation differs from published SMOG
This is worth stating plainly before anything else. Published SMOG is defined over a sample of exactly 30 sentences: ten from the start of a document, ten from the middle, ten from the end. You count the polysyllabic words in that sample, take the square root, and add a constant.
This tool does not ask you to sample. It counts polysyllables across whatever text you paste and normalises to a 30-sentence equivalent — the polysyllable count is scaled by 30 ÷ sentence count before the square root, then run through the same 1.043 and 3.1291 constants. On a document of reasonable length the result tracks the sampled figure closely, which is the point of the normalisation.
On very short text it does not, and cannot. Scaling a handful of sentences up to a 30-sentence equivalent multiplies whatever the polysyllable rate happens to be in those few sentences, so a paragraph with one unusually long word can read several grades harder than the document it came from. If you need a figure defensible against the published method, give it at least 30 sentences.
Why health communicators use it
SMOG is the standard in health literacy, and the reason is a calibration difference rather than a formula difference. Most readability formulas are fitted to roughly 50–75% comprehension of the passage. SMOG was fitted to 100% comprehension — full understanding, not the gist.
That makes SMOG scores run about one to two grades higher than Flesch–Kincaid on the same text, which looks like pessimism and is actually a different question being answered. When the cost of partial comprehension is someone taking a medication wrongly, the stricter target is the right one.
It also means you should not mix the two in a single report without saying which is which. A document at "grade 8 by Flesch–Kincaid" and "grade 10 by SMOG" has not got worse between two lines of a table.