It's hard to generalize from individual testing of specific units, but then again QA should provide some basic uniformity across units. Also, units may sound great in some configurations wiht some sources and some speakers. It's immensly complicated system to normalize for objective testing, as individual users report issues from their specific configuraitons of devices.
It depends which features are tested. That's why we need more comprehensive insight into various configurations, so that we could find out about dialogie bleeding in 7.2.4, etc. Someone needs to volunteer to list all possible configurations and issues/non-issues identified. Otherwise, we have informational chaos.