While the ability to store ultrasonics is something a higher sampling rate can do, that is not the only benefit to it. 44.1khz wasn't chosen because it was determined to be adequate, it was sort of a historical accident, mainly involving compatibility with the size of CD and video recording of the day. The antialiasing/reconstruction filters required for 44.1khz mangle transients and smear the temporal resolution. 44.1khz is not high enough to accurately capture or reproduce the impulse of things like a cymbal strike, which shows up on a spectrum analysis as ultrasonic information. Fourier transformation into the frequency domain only applies to sine waves, music is a complex wave, not pure sinewave tones, and the ear is more sensitive to temporal information than frequency.
Your post is filled with misinformation.
A 44.1KHz sampling rate is adequate for accurately reproducing any frequency up to 22KHz, which is beyond normal human hearing. For anyone over 40 years old, substantially beyond. Your statement about music being "complex waves", and FT being applicable only to pure sine waves is just plain incorrect.
You're also wrong about the Redbook standard being some kind of historical accident. (The Compact Disc specification was originally printed with a red cover, so it was called The Redbook.) The 44.1KHz sampling rate was chosen due to compatibility with the recordable bandwidth of video recorders of the day (late 1970s), because by leveraging video recording hardware Sony and Phillips (the two companies driving the Redbook standard) could get to market more quickly and cheaply, while still providing a 0-20KHz frequency response, which was the objective. 44.1KHz also provides a narrow guard band for the audio spectrum with the digital filtering necessary to remove aliasing for samples above one half the sampling frequency, which is 22.05KHz for CD audio.
The word width to represent amplitude (a sample is just a word containing the amplitude; the frequency is calculated in reference to a clock) was a point of contention between Sony and Phillips. Phillips thought a 14bit word width would be sufficient for CDs, since the resulting ~80db native dynamic range was about equal to the best dolby-enabled analog tape recorders of the day. Sony insisted on a 16bit word width to make CD audio truly superior to anything in analog, and perhaps to reset the playing field for the Phillips TDA-1540 14bit DAC that was already under development, and supposedly ahead of Sony's DAC effort. (Cynic that I am, I tend to believe the latter theory.)
Word width is where digital audio representation strategies do differ, because the greater the word width the more headroom an engineer gets during the recording process before digital clipping occurs. (Digital clipping (running out of amplitude bits) is catastrophic, and results in gross distortion, unlike analog clipping, which is soft and progressive.) So for recording engineers you would want, say, a 20-24bit word width for recording, and then for mastering (creating the final version of the recording for CD audio) you would cleverly truncate the extra bits in software, and reposition the amplitudes of the samples for maximum utilization of the 16bit CD words. But for audio playback, word widths greater than 16bits are just marketing nonsense, in the if-16-is-good-24-must-be-better category of foolishness.
As for aliasing filters "mangling" the audio band, that's bullshit. With modern digital filters and 8x or more over-sampling rates the reconstructed CD audio band is pretty much perfect down to 90db+ below the fundamental amplitudes of the frequencies, which are reproduced perfectly from 0-20KHz.
And the ear is more sensitive to temporal differences than frequency differences? Seriously? Who says that? MQA marketing?