This is downstream of how the models are built. A generator tuned hard on photographic data learns skin, light falloff and lens behaviour, and it gets stylised prompts subtly wrong β the linework is soft, the shading is photographic where it should be flat, and faces drift toward realism no matter what you ask for. A generator tuned on illustration learns line weight, colour blocking and the specific conventions of anime faces, and its photoreal attempts come out looking airbrushed.
Platforms know this, which is why several ship two separate models behind a style toggle rather than one model with a style prompt. Where you see a genuine toggle, output is usually good on both sides. Where style is just a word you type into the prompt, one side will be clearly weaker.
This is the single biggest thing an averaged image-quality score hides, and it is why we would tell someone to decide on style before looking at any number on this site.