Following up on my previous post, a few more points are worth making
regarding the scientific process.
First, regarding uncertainty. Earth's atmosphere and oceans do form a
more complicated system than the simple model I described. For
example, here's one way in which it is possible that temperatures
would not rise much in response to carbon dioxide impeding the outward
flow of heat. When temperatures go up initially, that means more
water vapor in the atmosphere. If that water vapor condenses into
clouds, the extra clouds could reflect enough sunlight back into space
to reduce the heating and make temperatures fall back to normal. This
mechanism would act like a thermostat keeping Earth's surface within a
narrow temperature range, and we wouldn't need to worry about keeping
our carbon emissions in check. So, if you heard Arrhenius's warming
prediction in 1896, you could easily say, "but there's a lot of
uncertainty in that prediction because we don't understand cloud
formation. Maybe there won't be that much warming. It's uncertain."
The point I want to make is that uncertainty cuts both ways. Water
vapor is itself a greenhouse gas, so if the extra vapor does not
condense into clouds, the greenhouse warming will be accelerated.
Yes, the prediction is uncertain....but that means that more extreme
outcomes, as well as less extreme outcomes, are possible.
If a little bit of warming produces clouds which shut down further
warming, we would call that a negative feedback loop; negative because
any change contains the seeds of its own reversal. If instead a bit
of warming creates water vapor which accelerates the warming, we would
call that a positive feedback loop; positive because a little movement
encourages further movement in the same direction. One reason climate
is complicated is that it is full of feedback loops, another example
being that reduced ice coverage causes more sunlight to be absorbed,
which reduces ice coverage further, etc. So what's the verdict on the
cloud formation? We still don't know; it may depend on how much
small-particle pollution we produce, because these small particles
provide the seeds for cloud condensation. But meanwhile, temperatures
keep rising. So while we puzzle over the details, let's not forget
the big picture: we keep making Earth's carbon-dioxide blanket thicker
and thicker.
Second point: I've repeatedly stressed the important of models in terms of
understanding a system. Models are great for exploring a variety of
scenarios, but is there anything we can say about climate that does
not depend on what model we adopt? Such model-independent statements
can be valuable anchors when we're not sure which model to adopt. I'd
like to focus particularly on a (more or less) model-independent
statement regarding sea levels. We can get rid of models and just
accumulate data regarding sea levels and carbon dioxide levels in the
past, and then we can simply ask, what is the typical sea level when
the carbon dioxide level is 400 parts per million, as we have now
caused it to be? (It's up from about 275 before the Industrial
Revolution.)
The answer is shocking: 24 meters, or 80 feet! Go ahead and play with
this interactive flood map to see what such a rise will do to your
state or country.
Now I have to give a few caveats. First, changes in carbon dioxide
concentration and sea levels occurred very slowly in the past.
Although we are pumping carbon dioxide in very quickly, it's quite
likely that it will be hundreds or even a few thousand years before
the effects of the carbon input are fully realized and sea levels rise
this much. Essentially no one is predicting these sea levels within
our children's lifetimes. But still....this will be a lot for our
great-great-grandchildren to deal with. And yes, there's uncertainty on this
prediction. Sea levels may rise less than this. But they may also rise more
than this.
Second caveat: this prediction is not entirely model-independent. To
be an extreme devil's advocate, if CO2 levels in the past were somehow
a natural effect of higher sea levels rather than a cause, then we could not
use past data to predict what would happen when we artificially increase
CO2 levels today. To be clear, I invoke that scenario not because I
believe it, but simply to highlight how an apparently
model-independent statement is often not entirely
model-independent. If all kinds of crazy models are allowed into the
discussion, then very few truly model-independent statements can be
made. But within the scope of "reasonable" models, we can say that
sea levels will rise by around 24 meters; we just don't know long
that will take. Predicting how long it will take requires a model!
If you are interested in further reading, start with this Skeptical
Science post, which summarizes this publication in an approachable
way. Skeptical Science, by the way, is a good resource for rebutting
common climate myths.
Thursday, February 27, 2014
Wednesday, February 26, 2014
Climate 101
Nice article today in the New York Times: Study Links Temperature to a
Peruvian Glacier’s Growth and Retreat. It's a good example of how news
about climate change could easily be misread as indicating more doubt
than there really is. The headline makes it sound as if the link between
glaciers and temperature is so tenuous that this is the first evidence of it,
and that it has been established for only one glacier. The truth is very
different, even though the headline and article are not wrong once you
understand the context. This post is aimed at helping teachers
and students with the basics, and then use that to parse the news.
Over a century ago, it was known that carbon dioxide impedes the flow
of heat (in the form of infrared light) from the Earth out into space,
while not impeding the flow of heat (mostly in the form of visible
light) from the Sun to the Earth. If not for this natural greenhouse
effect, Earth would be much colder. Teachers can demonstrate quite
directly that carbon dioxide impedes the flow of infrared light, but
many teachers may not have the right equipment. Here's a video
comparing the temperature rise of two bottles, one with elevated
levels of carbon dioxide and the other with standard air. And here's
a nice video using an infrared camera to show quite directly that
infrared light is largely blocked by carbon dioxide.
Around the same time (1896) Svante Arrhenius recognized that humans
were pumping ever more carbon dioxide into the atmosphere, and that
this would lead to warming. But "warming" sounded reasonably
beneficial, especially given Arrhenius's prediction that it would take
place slowly over thousands of years. Arrhenius did not account for
the large increase in population over the ensuing century, nor for the
large increase in per-capita use of fossil fuels (cars, airplanes,
etc). Worldwide, we now emit about 17 times the carbon dioxide emitted
in 1896, so change is coming much faster. And now we know that an
increase in temperature is not as beneficial as it may sound because
it can radically change weather patterns, which imposes large costs on
humans as well as on many species which cannot move and adapt rapidly
enough. Apart from that, Arrhenius deserves kudos for his prescience.
Yet if we heard this prediction in 1896 we would be justified in
expressing some skepticism. Earth's atmosphere and oceans (where most
of the excess heat is deposited) form a complicated system, and the
response of a complicated system to a simple input (more heat) may
well not be a simple result (higher temperature). But healthy
skepticism goes only so far; unless you have a better model, you have
to admit that the best model predicts warming. Just saying "it's a
complicated system" does not give you the right to reject all models.
In this case, you would have to figure out where the extra heat would
go without causing increased temperatures, and you would have to have
some evidence to motivate belief in that model.
Fast forward to 2014. Warming is here, and we've learned a lot about
climate models in the meantime. We did find complications (El Nino,
for one), but the simple model was reasonable in its overall
prediction. More heat does mean a higher temperature.
One way "climate skeptics" (I put the term in quotes because
oil-company funding leads to a kind of "skepticism" different from the
detached sort of skepticism we encourage in science) sow doubt about
this result is to suggest that the warming may be due to natural
cycles. There certainly are natural climate cycles, but rather than
treat them in detail here I want to make a bigger point about how
science works: When a model makes a prediction and the prediction
comes true, we should gain confidence in the model, and we should lose
confidence in models which made contrary predictions. Yes, it's
conceivable that the greenhouse model's prediction came true through
a fluke of natural cycles rather than accurately modeling how nature works
...but how much confidence would you put on that possibility?
A prediction is a powerful thing, so let's note the distinction between
a prediction and a retrodiction (or postdiction), which is when you make a
hypothesis after looking at the data. Using data (rather than laws of
physics or other guiding principles) to generate hypotheses is a fine
thing to do, but because "patterns" can randomly appear in data you
cannot confirm the hypothesis with the same data which generated it;
you must seek out new data. (Admittedly, even scientists sometimes
forget to apply this principle.) Climate skeptics can suggest
alternative causes for the warming after looking at the data, but we
should have much more confidence in the model which actually predicted
the data.
Now, the news: a reconstruction of the timeline of growth and
shrinkage of a Peruvian glacier shows that shrinkage is most highly
correlated with temperature and not with other factors such as
precipitation. You have to get halfway through the article to get the
background:
So, you may have started reading the article thinking that scientists
understood very little about glaciers if they were just now finding a
"link" between glacier shrinkage and temperature, but you now see
that a lot of important knowledge has already been established.
Newspaper articles are designed to tell you what's new first, so it's not
the writer's fault that this background was buried deep in the article.
Nevertheless, in practice many readers will just read the headline and
skim the first part of the article, thus missing this crucial background.
Teachers and students should be aware of this when reading science news.
But wait, there's more! The article goes on to explain how the details
of tropical glaciers are different from most glaciers (intense
sunlight can vaporize the ice directly, and the sunlight lasts
year-round) but that one group of scientists has studied the matter
and still concluded that temperature is the driving factor in
shrinking tropical glaciers. "But a second group believes that in some
circumstances, at least, a tropical glacier’s long-term fate may
reflect other factors. In particular, these scientists believe big
changes in precipitation can sometimes have more of a role than
temperature." In other words, this is a legitimate scientific dispute, but
it is about the details of a very specific type of glacier and has little or
nothing to do with overall concerns about glaciers (or sea ice)
melting worldwide, much less about the reality of climate change. Yet
someone who wants to sow doubt about climate change can point to this
and say "scientists don't really understand why glaciers melt" and people
who don't read the article carefully may well be snookered by that.
Please make sure you (and, if you are a teacher, your students) don't
get snookered.
My next post discusses two more aspects of the nature of science---uncertainty
and model-independent statements---in the context of climate.
Peruvian Glacier’s Growth and Retreat. It's a good example of how news
about climate change could easily be misread as indicating more doubt
than there really is. The headline makes it sound as if the link between
glaciers and temperature is so tenuous that this is the first evidence of it,
and that it has been established for only one glacier. The truth is very
different, even though the headline and article are not wrong once you
understand the context. This post is aimed at helping teachers
and students with the basics, and then use that to parse the news.
Over a century ago, it was known that carbon dioxide impedes the flow
of heat (in the form of infrared light) from the Earth out into space,
while not impeding the flow of heat (mostly in the form of visible
light) from the Sun to the Earth. If not for this natural greenhouse
effect, Earth would be much colder. Teachers can demonstrate quite
directly that carbon dioxide impedes the flow of infrared light, but
many teachers may not have the right equipment. Here's a video
comparing the temperature rise of two bottles, one with elevated
levels of carbon dioxide and the other with standard air. And here's
a nice video using an infrared camera to show quite directly that
infrared light is largely blocked by carbon dioxide.
Around the same time (1896) Svante Arrhenius recognized that humans
were pumping ever more carbon dioxide into the atmosphere, and that
this would lead to warming. But "warming" sounded reasonably
beneficial, especially given Arrhenius's prediction that it would take
place slowly over thousands of years. Arrhenius did not account for
the large increase in population over the ensuing century, nor for the
large increase in per-capita use of fossil fuels (cars, airplanes,
etc). Worldwide, we now emit about 17 times the carbon dioxide emitted
in 1896, so change is coming much faster. And now we know that an
increase in temperature is not as beneficial as it may sound because
it can radically change weather patterns, which imposes large costs on
humans as well as on many species which cannot move and adapt rapidly
enough. Apart from that, Arrhenius deserves kudos for his prescience.
Yet if we heard this prediction in 1896 we would be justified in
expressing some skepticism. Earth's atmosphere and oceans (where most
of the excess heat is deposited) form a complicated system, and the
response of a complicated system to a simple input (more heat) may
well not be a simple result (higher temperature). But healthy
skepticism goes only so far; unless you have a better model, you have
to admit that the best model predicts warming. Just saying "it's a
complicated system" does not give you the right to reject all models.
In this case, you would have to figure out where the extra heat would
go without causing increased temperatures, and you would have to have
some evidence to motivate belief in that model.
Fast forward to 2014. Warming is here, and we've learned a lot about
climate models in the meantime. We did find complications (El Nino,
for one), but the simple model was reasonable in its overall
prediction. More heat does mean a higher temperature.
One way "climate skeptics" (I put the term in quotes because
oil-company funding leads to a kind of "skepticism" different from the
detached sort of skepticism we encourage in science) sow doubt about
this result is to suggest that the warming may be due to natural
cycles. There certainly are natural climate cycles, but rather than
treat them in detail here I want to make a bigger point about how
science works: When a model makes a prediction and the prediction
comes true, we should gain confidence in the model, and we should lose
confidence in models which made contrary predictions. Yes, it's
conceivable that the greenhouse model's prediction came true through
a fluke of natural cycles rather than accurately modeling how nature works
...but how much confidence would you put on that possibility?
A prediction is a powerful thing, so let's note the distinction between
a prediction and a retrodiction (or postdiction), which is when you make a
hypothesis after looking at the data. Using data (rather than laws of
physics or other guiding principles) to generate hypotheses is a fine
thing to do, but because "patterns" can randomly appear in data you
cannot confirm the hypothesis with the same data which generated it;
you must seek out new data. (Admittedly, even scientists sometimes
forget to apply this principle.) Climate skeptics can suggest
alternative causes for the warming after looking at the data, but we
should have much more confidence in the model which actually predicted
the data.
Now, the news: a reconstruction of the timeline of growth and
shrinkage of a Peruvian glacier shows that shrinkage is most highly
correlated with temperature and not with other factors such as
precipitation. You have to get halfway through the article to get the
background:
land ice is melting virtually everywhere on the planet...the pace seems to have accelerated substantially in recent decades as human emissions have begun to overwhelm the natural cycles. In the middle and high latitudes, from Switzerland to Alaska, a half-century of careful glaciology has established that temperature is the main factor controlling the growth and recession of glaciers. But the picture has been murkier in the tropics. There, too, glaciers are retreating, but scientists have had more trouble sorting out exactly why.
So, you may have started reading the article thinking that scientists
understood very little about glaciers if they were just now finding a
"link" between glacier shrinkage and temperature, but you now see
that a lot of important knowledge has already been established.
Newspaper articles are designed to tell you what's new first, so it's not
the writer's fault that this background was buried deep in the article.
Nevertheless, in practice many readers will just read the headline and
skim the first part of the article, thus missing this crucial background.
Teachers and students should be aware of this when reading science news.
But wait, there's more! The article goes on to explain how the details
of tropical glaciers are different from most glaciers (intense
sunlight can vaporize the ice directly, and the sunlight lasts
year-round) but that one group of scientists has studied the matter
and still concluded that temperature is the driving factor in
shrinking tropical glaciers. "But a second group believes that in some
circumstances, at least, a tropical glacier’s long-term fate may
reflect other factors. In particular, these scientists believe big
changes in precipitation can sometimes have more of a role than
temperature." In other words, this is a legitimate scientific dispute, but
it is about the details of a very specific type of glacier and has little or
nothing to do with overall concerns about glaciers (or sea ice)
melting worldwide, much less about the reality of climate change. Yet
someone who wants to sow doubt about climate change can point to this
and say "scientists don't really understand why glaciers melt" and people
who don't read the article carefully may well be snookered by that.
Please make sure you (and, if you are a teacher, your students) don't
get snookered.
My next post discusses two more aspects of the nature of science---uncertainty
and model-independent statements---in the context of climate.
Friday, February 21, 2014
One Percenters
We've been bombarded all winter with stories of cold and snowy weather in the eastern US, but the news was just released that January 2014 was the fourth-warmest January on record. How can this be? The eastern US covers less than 1% of the Earth's area, so (as this essay nicely puts it) "if the whole country somehow froze solid one January, that would not move the needle on global temperatures much at all." That essay is worth reading because it goes on to explain how subjectively people do perceive global warming: something as unrelated to global warming as being in a cold room does have an influence on the opinions voiced in a survey. Educators should be aware of this, and actively work on making students think objectively and use data.
Wednesday, January 29, 2014
Mostly Harmless
In the Hitchhiker's Guide to the Galaxy, "mostly harmless" is the
Encyclopedia Galactica's assessment of Earth (which is not important
enough to merit a longer entry). This made me think that looking at
the solar system through alien's eyes might help students learn about
it. I conducted Science in the River City workshop for earth science
teachers based on this idea, and this is a list of resources for such
teachers.
First, I highlighted a graphing activity I had done with elementary
kids; that experienced is described in great detail here. (Feel free
to download and copy the graph.) I extended the activity to graphing the
surface temperatures of the planets as a function of distance from the
Sun, which led to the greenhouse effect discussion below, but now it
occurs to me that a great way to extend this activity would be to jigsaw
it: assign one group of students to graph size vs distance from the Sun,
another to graph temperature vs distance from the Sun, another to graph
density vs distance from the Sun, etc, and then the groups come together
to think about what it all implies for the formation of the solar system.
Encyclopedia Galactica's assessment of Earth (which is not important
enough to merit a longer entry). This made me think that looking at
the solar system through alien's eyes might help students learn about
it. I conducted Science in the River City workshop for earth science
teachers based on this idea, and this is a list of resources for such
teachers.
First, I highlighted a graphing activity I had done with elementary
kids; that experienced is described in great detail here. (Feel free
to download and copy the graph.) I extended the activity to graphing the
surface temperatures of the planets as a function of distance from the
Sun, which led to the greenhouse effect discussion below, but now it
occurs to me that a great way to extend this activity would be to jigsaw
it: assign one group of students to graph size vs distance from the Sun,
another to graph temperature vs distance from the Sun, another to graph
density vs distance from the Sun, etc, and then the groups come together
to think about what it all implies for the formation of the solar system.
Second, when discussing the formation of the solar system and
describing how small grains of dust started to stick together, I
wanted to show a video clip but had some technical difficulties. Here
is the link; start at 3 minutes into the video and go for 2.5 minutes.
(If you have time, the whole episode is worth watching. It's from the
How the Earth Was Made series, which has some really nice
visualizations and is constructed around evidence, which is a key
feature missing from many science documentaries. It tells science
like the detective story it is. That's generally a good thing, but in
this case the implication that this particular astronaut doing this
particular demonstration singlehandedly saved the theory is a bit of
an exaggeration.)
Pocket solar system: https://nightsky.jpl.nasa.gov/download-view.cfm?Doc_ID=392
Peppercorn Earth: http://www.noao.edu/education/peppercorn/pcmain.html
Extrasolar planets: http://exoplanets.org/ has the most up-to-date
info. Even better, they have built-in graphing tools so you and
your students can easily explore the data.
Earth's surface temperature: I got my plot from the most authoritative
source for modern temperatures, NASA's Goddard Institute for Space Studies.
This link only scratches the surface of climate change data because it deals
with modern temperature measurements (as opposed to long-ago temperatures
inferred from ice cores etc) but as the greenhouse effect was not the focus of
the workshop I won't try to compile a list of links here. (For those
not attending the workshop: we graphed planets' surface temperatures
vs distance from the Sun, and we saw the general pattern that farther
from Sun equals colder, but we also saw that Venus is a real outlier
from this pattern. That's because Venus has had a runaway greenhouse
effect. Earth also has a natural greenhouse effect which keeps us
from being frozen, but which is now being augmented by a manmade
greenhouse effect. I did tell the teachers that Earth has a "carbon
cycle" which will absorb the extra carbon dioxide through the oceans
into rocks, but I forgot to mention that it will take hundreds of
thousands of years; I didn't mean to imply that humans can carry on
regardless. Venus' greenhouse effect is "runaway" because its
carbon cycle shut down when its oceans boiled.)
Finally, a few links I didn't get time to show but which will help you
appreciate the size of the universe (and the sizes of things in it):
the classic Powers of Ten video and an interactive tool.
Thursday, January 2, 2014
One Plus z
This marks the launch of a new series of posts, aimed at astronomy and physics majors. In the course of my teaching I've noticed a few topics---such as propagation of errors and reduced mass---which seem to fall through the cracks between classes. Students hear a bit about reduced mass in more than one class, but never seem to get a satisfying explanation in any one class. Their lab instructor taught them how to propagate errors but never made them think about why. And so on. This first post is much more specific---how to think about redshifts and velocity dispersions in cosmology---but fits the bill because it seems to fall through the cracks between textbooks. Practitioners know that "you need to divide by 1+z" but documentation of this is hard to come by. So here we go.
In cosmology, we often want to measure the rest-frame velocity dispersion of a galaxy cluster, but what we actually measure is the redshift dispersion. How are they related? Redshift z is defined in terms of emitted and observed wavelengths as
This means that 1+z is a stretching factor; it is the ratio of observed to emitted wavelengths. So you will see the combination 1+z over and over, rather than z by itself. Get used to thinking in terms of 1+z!
The Doppler shift formula tells us the wavelength stretching factor in terms of velocity:
You will often see this called the relativistic Doppler formula, as opposed to the simpler low-velocity approximation used in many situations. But I suggest thinking of this as the Doppler formula because high velocities are common in astrophysics, and this correct version is simple enough to memorize. Habitually using the low-velocity approximation can get you in trouble.
The Doppler formula can be inverted to obtain
Now imagine two galaxies, one at rest1 in the cluster frame (with velocity v1 in our frame) and a second moving with some velocity v21 relative to the cluster which implies some velocity v2 in our frame. According to the Einstein velocity addition law,
Substituting the inverted Doppler formula into this, we obtain a complicated-looking expression for v21/c:
which we can simplify in a few steps:
Because of my poor equation formatting, I have to remind you here that this is an expression for v21/c, where v21 represents a velocity in the cluster frame rather than in our frame. This gets us close to our goal because we want to know the velocity dispersion in the cluster frame. But this is as far as we can go without an approximation. A useful approximation in this context is that
so define
and eliminate z2 using
:
Taylor expanding this about
we obtain
This is true for any small redshift difference, so it must be true if delta represents the redshift dispersion of the cluster (thus making v21 represent the velocity dispersion of the cluster). Therefore
However, there is a much more elegant way to derive the same result. Imagine a hypothetical observer on the first galaxy. Because of the definition of 1+z as the ratio of wavelengths, it must be true that 1+z2 = (1+z1)(1+z21) where z21 is the redshift of galaxy 2 as seen by galaxy 1 (z1 and z2 are, as before, redshifts seen by us). Therefore
Again we use an approximation:
so that we can use the low-velocity approximation for the Doppler shift,
.
Therefore
which is the same result as before. We don't actually need special-relativistic reasoning if we simply use the definition of redshift to isolate the one nonrelativistic velocity in the problem.
We can better expose the equivalence of these two approaches by taking the idea of daisy-chaining wavelength ratios and applying it directly to the Doppler law:
This just says that galaxy 2's wavelength ratio ("ratio'' here is relative to a laboratory standard) observed by us is its wavelength ratio observed by galaxy 1, times galaxy 1's wavelength ratio observed by us. In a few lines of algebra, you can show that the above expression leads directly to the Einstein velocity addition law. The addition law can be derived in more than one way, but to me this is the most intuitive way. Thus, daisy-chaining Doppler factors and using the velocity addition law are not contrasting approaches; they are actually the same thing.
Exercise for the reader: show that the expression above does indeed lead to the Einstein velocity addition law.
Footnotes:
1 I specify "at rest" here only so that later it will be easy to think of this galaxy’s redshift as the mean cluster redshift.
In cosmology, we often want to measure the rest-frame velocity dispersion of a galaxy cluster, but what we actually measure is the redshift dispersion. How are they related? Redshift z is defined in terms of emitted and observed wavelengths as
This means that 1+z is a stretching factor; it is the ratio of observed to emitted wavelengths. So you will see the combination 1+z over and over, rather than z by itself. Get used to thinking in terms of 1+z!
The Doppler shift formula tells us the wavelength stretching factor in terms of velocity:
You will often see this called the relativistic Doppler formula, as opposed to the simpler low-velocity approximation used in many situations. But I suggest thinking of this as the Doppler formula because high velocities are common in astrophysics, and this correct version is simple enough to memorize. Habitually using the low-velocity approximation can get you in trouble.
The Doppler formula can be inverted to obtain
Now imagine two galaxies, one at rest1 in the cluster frame (with velocity v1 in our frame) and a second moving with some velocity v21 relative to the cluster which implies some velocity v2 in our frame. According to the Einstein velocity addition law,
Substituting the inverted Doppler formula into this, we obtain a complicated-looking expression for v21/c:
which we can simplify in a few steps:
Because of my poor equation formatting, I have to remind you here that this is an expression for v21/c, where v21 represents a velocity in the cluster frame rather than in our frame. This gets us close to our goal because we want to know the velocity dispersion in the cluster frame. But this is as far as we can go without an approximation. A useful approximation in this context is that
so define
and eliminate z2 using
:
Taylor expanding this about
we obtainThis is true for any small redshift difference, so it must be true if delta represents the redshift dispersion of the cluster (thus making v21 represent the velocity dispersion of the cluster). Therefore
However, there is a much more elegant way to derive the same result. Imagine a hypothetical observer on the first galaxy. Because of the definition of 1+z as the ratio of wavelengths, it must be true that 1+z2 = (1+z1)(1+z21) where z21 is the redshift of galaxy 2 as seen by galaxy 1 (z1 and z2 are, as before, redshifts seen by us). Therefore
Again we use an approximation:
so that we can use the low-velocity approximation for the Doppler shift,
.
Thereforewhich is the same result as before. We don't actually need special-relativistic reasoning if we simply use the definition of redshift to isolate the one nonrelativistic velocity in the problem.
We can better expose the equivalence of these two approaches by taking the idea of daisy-chaining wavelength ratios and applying it directly to the Doppler law:
This just says that galaxy 2's wavelength ratio ("ratio'' here is relative to a laboratory standard) observed by us is its wavelength ratio observed by galaxy 1, times galaxy 1's wavelength ratio observed by us. In a few lines of algebra, you can show that the above expression leads directly to the Einstein velocity addition law. The addition law can be derived in more than one way, but to me this is the most intuitive way. Thus, daisy-chaining Doppler factors and using the velocity addition law are not contrasting approaches; they are actually the same thing.
Exercise for the reader: show that the expression above does indeed lead to the Einstein velocity addition law.
Footnotes:
1 I specify "at rest" here only so that later it will be easy to think of this galaxy’s redshift as the mean cluster redshift.
Friday, December 20, 2013
Shedding light on missing mass
An important part of science---of life, I would argue---is making inferences from data. (Contrary to some popular perception, it is not all that we do, but it is a large part.) This process is a lot more interesting than many people think, and this is best conveyed by stories where it went horribly wrong. I just published a paper rebutting a paper in which it went wrong; not horribly so, but to make this post understandable I will weave a few of these horrible stories into my tale.
The paper I rebutted claimed that one method (called weak gravitational lensing) of measuring the mass of a certain galaxy cluster gave an answer too low compared to the answers obtained through two other methods, and therefore the lensing method itself was suspect. The context is that astronomers find it very difficult to measure the mass of anything, since we are so far away. If the cluster is not changing over time, we can relate the velocities of the galaxies in the cluster to its mass (called the dynamical method) and we can also relate the cluster's X-ray emission to its mass. But that's a big if, and we would like a method which does not depend on this assumption. Lensing is such a method; it has weaknesses too, but I don't want to get too deeply into that here. The central question in this paper is really simple and applies to many situations: when numbers seemingly disagree, how do we characterize the strength of disagreement given that there is some uncertainty associated with each number?
The original paper made a model of the cluster using the X-ray method, and simulated weak lensing measurements of this model to see how often the simulated measurements gave answers as low as the actual weak lensing measurements. This is a great technique; it gives us what's called a p-value. By tentatively assuming that weak lensing is as effective as the X-ray method---the "null hypothesis"---we will see how often the inherent uncertainties in weak lensing would just randomly give us an answer as low as we got in real life. If the answer is "never" then we can state that our null hypothesis is wrong and weak lensing is not as effective as the X-ray method. More quantitatively, if the answer is "in 1 out of every 100 experiments" we would say p=0.01, which has the naive interpretation of "99% confidence that the null hypothesis is rejected." (One of the reasons it's naive is that if you tested, say, 100 different true hypotheses, you would still expect one to randomly come out with p=0.01. So the true interpretation is more nuanced. I will develop this further below.)
Now, what if this method gives you p=0.1 or so? You can't really reject the null hypothesis unless you have stronger proof than that, so you may go out and take more data, do more experiments, etc, to get the stronger proof. If you do so, make sure that the new experiments are independent of the original one. For example, if you want to prove that tall people are better basketball players than short people, the null hypothesis would be that they are the same and you might record the score from a scrimmage in which a tall person plays against a short one. If the tall person comes out slightly ahead, you will not have strong proof that the tall person is better, so you might replay the scrimmage. But if you play the same two people against each other, you can never prove that tall is better; the most you might prove is that player A is better than player B. To make the trials independent, you have to play a different tall person against a different short person. In more general terms, if you're trying to get an idea of the natural variation or "noise" in your measurement, you have to repeat the measurement in a way that actually incorporates those variations. What this paper did was equivalent to failing to recognize the nonindependence of identical triplet weak lensing players. They ran three scrimmages between an X-ray player and each of these three weak lensing players, mistakenly yielding a strong conclusion about X-ray vs weak lensing.
This idea of independence---and recognizing nonindependence even when it's subtle---is really important. Ben Goldacre in his book Bad Science relates the story of a woman suspected of murder because two of her kids died of sudden infant death syndrome. The chance of one baby dying of SIDS was stated as 1 in 8543. Prosecutors assumed that the chance of a second child dying of SIDS (over the course of years, not in the same incident) was independent of the chance of the first child dying of SIDS, so we can multiply probabilities and come up with a 1 in 73,000,000 chance of two babies dying of SIDS; so unlikely that we might suspect murder. But they're not independent. If SIDS has anything to do with genes or environment then they can't be independent, because the babies have the same parents and the same house. Given the shared genes and environment, the second baby's chance of SIDS may actually be quite high. In that case, we have no reason to suspect murder. The prosecutors vastly overstated the statistical case for murder by failing to recognize the non-independence. (That's not the only mistake the prosecutors made. I highly recommend Goldacre's book.)
A second mistake the authors of the weak lensing paper made was multiplying the p-values from the three experiments to obtain an overall p-value. Many people, even scientists, fall into the following trap: Say Experiment A gives p=0.10 and you interpret that as only a 10% chance that the null hypothesis is correct. Now independent Experiment B gives p=0.08, which you interpret as only an 8% chance that the null hypothesis is correct. It is natural to think that the experiments together imply only 8% of a 10% chance of the null hypothesis being correct, or p=0.008. But it's wrong! You have vastly underestimated the chance of the null hypothesis being correct, just as the paper we rebutted vastly underestimated the chance that the weak lensing measurements were actually consistent with the dynamical and X-ray measurements. Even if the experiments are independent, you should not multiply the p-values.
Here's an easy way to confirm that the above procedure is wrong: following an equivalent procedure you could also interpret p=0.10 as a 90% chance that the null hypothesis is incorrect and p=0.08 as a 92% chance that the null hypothesis is incorrect. Multiplying them, we would get an 82.8% chance that the null hypothesis is incorrect. But the same process in the previous paragraph yielded an 0.8% chance that it's correct, which doesn't match the 82.8% chance that it's incorrect. So something's wrong with the process! And it gets more wrong as you do more experiments and more multiplications. If we do 100 independent experiments and follow this line of reasoning, we will come up with a vanishingly small chance of the null hypothesis being correct, and a vanishingly small chance of the null hypothesis being incorrect, regardless of the specific p-values, because they are always less than one. You will rule out things which do not deserve to be ruled out. Goldacre gives the horrifying example of a nurse who was suspected of murdering patients and convicted largely on the basis of faulty statistics but was eventually freed. I'm going to make up the following numbers for simplicity. Let's say that some number of patients died while she was working, such that there was only a 10% chance that that would have happened randomly. So you start poking around, and find that at the previous hospital where she worked, there was only a 50% chance of that large a number of patients dying, and at the hospital before that only a 70% chance, and at the hospital before that only a 30% chance, etc. Multiplying all these together gives a really small chance that all these things occurred randomly. But you know by now that multiplying these is wrong.
Why is it wrong? Probabilities can be multiplied (1/2 chance of heads in each coin toss means 1/4 chance of two heads in two coin tosses), but despite its name, the p-value is not a simple probability that the null hypothesis is true. It's a measure of consistency which is constructed so that for any one experiment we interpret all but the lowest p-values as being consistent with the null hypothesis. Therefore p=0.5, say, is perfectly consistent with the null hypothesis; it does not mean a 50% chance of it being true or false. In the hypothetical example above, p=0.50 actually means that an average number of patients died on shift (deaths on random shifts rose to [at least] that level 50% of the time) and p=0.70 actually means that fewer than average patients died on shift (deaths on random shifts rose to [at least] that level 70% of the time). A correct way to combine p-values for independent experiments is Fisher's method. Had the paper we rebutted used that method, they would have seen that the dynamical and weak lensing measurements were entirely consistent, even without correcting the error regarding nonindependent trials. Correcting both errors makes it even more clear.
This stuff is complicated and it's easy to go wrong. Happily, many incorrect inferences in science are caught relatively quickly because so many scientists have so much practice in this kind of analysis. But Goldacre's book is an eye-opener. Things don't always work out so well so quickly.
The paper I rebutted claimed that one method (called weak gravitational lensing) of measuring the mass of a certain galaxy cluster gave an answer too low compared to the answers obtained through two other methods, and therefore the lensing method itself was suspect. The context is that astronomers find it very difficult to measure the mass of anything, since we are so far away. If the cluster is not changing over time, we can relate the velocities of the galaxies in the cluster to its mass (called the dynamical method) and we can also relate the cluster's X-ray emission to its mass. But that's a big if, and we would like a method which does not depend on this assumption. Lensing is such a method; it has weaknesses too, but I don't want to get too deeply into that here. The central question in this paper is really simple and applies to many situations: when numbers seemingly disagree, how do we characterize the strength of disagreement given that there is some uncertainty associated with each number?
The original paper made a model of the cluster using the X-ray method, and simulated weak lensing measurements of this model to see how often the simulated measurements gave answers as low as the actual weak lensing measurements. This is a great technique; it gives us what's called a p-value. By tentatively assuming that weak lensing is as effective as the X-ray method---the "null hypothesis"---we will see how often the inherent uncertainties in weak lensing would just randomly give us an answer as low as we got in real life. If the answer is "never" then we can state that our null hypothesis is wrong and weak lensing is not as effective as the X-ray method. More quantitatively, if the answer is "in 1 out of every 100 experiments" we would say p=0.01, which has the naive interpretation of "99% confidence that the null hypothesis is rejected." (One of the reasons it's naive is that if you tested, say, 100 different true hypotheses, you would still expect one to randomly come out with p=0.01. So the true interpretation is more nuanced. I will develop this further below.)
Now, what if this method gives you p=0.1 or so? You can't really reject the null hypothesis unless you have stronger proof than that, so you may go out and take more data, do more experiments, etc, to get the stronger proof. If you do so, make sure that the new experiments are independent of the original one. For example, if you want to prove that tall people are better basketball players than short people, the null hypothesis would be that they are the same and you might record the score from a scrimmage in which a tall person plays against a short one. If the tall person comes out slightly ahead, you will not have strong proof that the tall person is better, so you might replay the scrimmage. But if you play the same two people against each other, you can never prove that tall is better; the most you might prove is that player A is better than player B. To make the trials independent, you have to play a different tall person against a different short person. In more general terms, if you're trying to get an idea of the natural variation or "noise" in your measurement, you have to repeat the measurement in a way that actually incorporates those variations. What this paper did was equivalent to failing to recognize the nonindependence of identical triplet weak lensing players. They ran three scrimmages between an X-ray player and each of these three weak lensing players, mistakenly yielding a strong conclusion about X-ray vs weak lensing.
This idea of independence---and recognizing nonindependence even when it's subtle---is really important. Ben Goldacre in his book Bad Science relates the story of a woman suspected of murder because two of her kids died of sudden infant death syndrome. The chance of one baby dying of SIDS was stated as 1 in 8543. Prosecutors assumed that the chance of a second child dying of SIDS (over the course of years, not in the same incident) was independent of the chance of the first child dying of SIDS, so we can multiply probabilities and come up with a 1 in 73,000,000 chance of two babies dying of SIDS; so unlikely that we might suspect murder. But they're not independent. If SIDS has anything to do with genes or environment then they can't be independent, because the babies have the same parents and the same house. Given the shared genes and environment, the second baby's chance of SIDS may actually be quite high. In that case, we have no reason to suspect murder. The prosecutors vastly overstated the statistical case for murder by failing to recognize the non-independence. (That's not the only mistake the prosecutors made. I highly recommend Goldacre's book.)
A second mistake the authors of the weak lensing paper made was multiplying the p-values from the three experiments to obtain an overall p-value. Many people, even scientists, fall into the following trap: Say Experiment A gives p=0.10 and you interpret that as only a 10% chance that the null hypothesis is correct. Now independent Experiment B gives p=0.08, which you interpret as only an 8% chance that the null hypothesis is correct. It is natural to think that the experiments together imply only 8% of a 10% chance of the null hypothesis being correct, or p=0.008. But it's wrong! You have vastly underestimated the chance of the null hypothesis being correct, just as the paper we rebutted vastly underestimated the chance that the weak lensing measurements were actually consistent with the dynamical and X-ray measurements. Even if the experiments are independent, you should not multiply the p-values.
Here's an easy way to confirm that the above procedure is wrong: following an equivalent procedure you could also interpret p=0.10 as a 90% chance that the null hypothesis is incorrect and p=0.08 as a 92% chance that the null hypothesis is incorrect. Multiplying them, we would get an 82.8% chance that the null hypothesis is incorrect. But the same process in the previous paragraph yielded an 0.8% chance that it's correct, which doesn't match the 82.8% chance that it's incorrect. So something's wrong with the process! And it gets more wrong as you do more experiments and more multiplications. If we do 100 independent experiments and follow this line of reasoning, we will come up with a vanishingly small chance of the null hypothesis being correct, and a vanishingly small chance of the null hypothesis being incorrect, regardless of the specific p-values, because they are always less than one. You will rule out things which do not deserve to be ruled out. Goldacre gives the horrifying example of a nurse who was suspected of murdering patients and convicted largely on the basis of faulty statistics but was eventually freed. I'm going to make up the following numbers for simplicity. Let's say that some number of patients died while she was working, such that there was only a 10% chance that that would have happened randomly. So you start poking around, and find that at the previous hospital where she worked, there was only a 50% chance of that large a number of patients dying, and at the hospital before that only a 70% chance, and at the hospital before that only a 30% chance, etc. Multiplying all these together gives a really small chance that all these things occurred randomly. But you know by now that multiplying these is wrong.
Why is it wrong? Probabilities can be multiplied (1/2 chance of heads in each coin toss means 1/4 chance of two heads in two coin tosses), but despite its name, the p-value is not a simple probability that the null hypothesis is true. It's a measure of consistency which is constructed so that for any one experiment we interpret all but the lowest p-values as being consistent with the null hypothesis. Therefore p=0.5, say, is perfectly consistent with the null hypothesis; it does not mean a 50% chance of it being true or false. In the hypothetical example above, p=0.50 actually means that an average number of patients died on shift (deaths on random shifts rose to [at least] that level 50% of the time) and p=0.70 actually means that fewer than average patients died on shift (deaths on random shifts rose to [at least] that level 70% of the time). A correct way to combine p-values for independent experiments is Fisher's method. Had the paper we rebutted used that method, they would have seen that the dynamical and weak lensing measurements were entirely consistent, even without correcting the error regarding nonindependent trials. Correcting both errors makes it even more clear.
This stuff is complicated and it's easy to go wrong. Happily, many incorrect inferences in science are caught relatively quickly because so many scientists have so much practice in this kind of analysis. But Goldacre's book is an eye-opener. Things don't always work out so well so quickly.
Monday, September 2, 2013
"Just" a Theory?
A recently published letter to the New York Times reminds us that relativity is "just a theory" and so is the Big Bang. Scientists and science educators need to set the record straight on this "just a theory" meme any time we get a chance to discuss science with kids and grown-up nonscientists. So here's my shot at it.
A good analogy is to think of facts as being like bricks: solid and dependable, but one or a few bricks are not very useful by themselves ("an electron passed through my detector at 11:58:32.01" or "the high temperature in Davis, CA on September 1, 2013 was 96 F"). Only when we assemble lots (lots) of bricks into a coherent structure do we get the benefits of having a building (the theory of relativity, or a climate model). Not only is an isolated brick rather useless, but the building can easily survive the removal of a few bricks here and there. A good theory integrates millions or billions of observations into a coherent whole. Calling relativity "just a theory" is like calling the Great Wall of China "just a fence," the Panama Canal "just a ditch," or the Golden Gate Bridge "just a road."
There's a reason that calling the Great Wall of China "just a fence" sounds more outrageous than calling relativity "just a theory"---I used the word fence which connotes something less important than a wall. There's a rich vocabulary to describe to describe barriers: from weak to strong we might use tape, rope, cordon, railing, fence, and wall. But most people don't use a similarly rich vocabulary to describe levels of sophistication of mental models. From weak to strong I might suggest educated guess, working hypothesis, model, and theory, but most people in practice indiscriminately use the word theory for any of these. So it's our duty as scientists to make clear that well-accepted scientific theories integrate an incredible range of observations into a structure which is so coherent that it is difficult to imagine all those pieces fitting into any other structure. Maybe a better analogy to calling relativity "just a theory" is calling an assembled jigsaw puzzle "just one way to fit the pieces together."
Gotcha, the just-a-theory crowd says, by making that analogy you are showing that you are rigid in your thinking and unwilling to accept alternative explanations. Nonsense. Scientists are constantly trying to prove accepted theories wrong. Anyone who succeeds in disproving relativity, the Big Bang, or evolution will win a Nobel Prize and eternal fame, so we'd be happy to do so. But we know from experience that the most likely explanation for an isolated fact that seems to contradict relativity, the Big Bang, or evolution is that the fact itself was taken out of context or is not being properly interpreted, rather than that an extremely well-tested theory is wrong.
This doesn't mean that we will twist any fact to make it fit into our well-accepted theories. It does mean that surprising facts may end up extending the theory rather than replacing it. For example, Newton's theory of gravity explains a ton of observations about the motions of the planets and stars, but in a few extreme circumstances (such as very close to the Sun) it doesn't predict exactly what is observed. Einstein developed a theory of gravity (general relativity) which does correctly predict these situations. Einstein's theory is more complicated than Newton's, but in most situations the complicated parts of Einstein's theory have very little quantitative effect so we can simplify it a great deal and in those cases it turns out to be identical to....Newton's theory! This almost had to be the case, because Newton's theory accounted so well for so many observations that it would be hard to imagine that it was wrong rather than incomplete.
This example shows that a small number of facts can be critically important and that scientists do pay attention to facts which don't fit the theory. But we don't modify or overturn theories willy-nilly. When the planet Uranus didn't move exactly as Newton's theory predicted, modifications of the theory were considered but so was the possibility that some mass other than the Sun and the known planets was pulling on Uranus, and that led to the discovery of Neptune. If we rejected well-established theories at the first hint of any discrepancy with new observations, we would be giving undue weight to the new observations and too little weight to the vast range of previous observations explained by the theory. If you want to overthrow a theory because some new observation seems to contradict it, then give us a better theory which explains the new observation while still fitting the previous observations just as well as the old theory. That latter part seems to be conveniently forgotten by people who want to reject well-established theories.
A closely parallel situation is that of criminal investigators and prosecutors who present their "theory of the crime" to a jury. ("Model of the crime" would better fit my vocabulary hierarchy, but this is the word actually used.) A lot of facts may be introduced into evidence ("a car with the suspect's license plate was recorded crossing the Tappan Zee Bridge at 2:20am on August 31"), but by themselves they don't mean anything important. A good theory of the crime provides a coherent explanation of so many different facts that the jury is forced to conclude that it is true beyond a reasonable doubt. If you want to call it "just a theory" then offer us a different theory which fits the facts just as well. The defense is given sufficient time and strong motivation to offer a good alternative theory, so failure to present one is damning.
A good analogy is to think of facts as being like bricks: solid and dependable, but one or a few bricks are not very useful by themselves ("an electron passed through my detector at 11:58:32.01" or "the high temperature in Davis, CA on September 1, 2013 was 96 F"). Only when we assemble lots (lots) of bricks into a coherent structure do we get the benefits of having a building (the theory of relativity, or a climate model). Not only is an isolated brick rather useless, but the building can easily survive the removal of a few bricks here and there. A good theory integrates millions or billions of observations into a coherent whole. Calling relativity "just a theory" is like calling the Great Wall of China "just a fence," the Panama Canal "just a ditch," or the Golden Gate Bridge "just a road."
There's a reason that calling the Great Wall of China "just a fence" sounds more outrageous than calling relativity "just a theory"---I used the word fence which connotes something less important than a wall. There's a rich vocabulary to describe to describe barriers: from weak to strong we might use tape, rope, cordon, railing, fence, and wall. But most people don't use a similarly rich vocabulary to describe levels of sophistication of mental models. From weak to strong I might suggest educated guess, working hypothesis, model, and theory, but most people in practice indiscriminately use the word theory for any of these. So it's our duty as scientists to make clear that well-accepted scientific theories integrate an incredible range of observations into a structure which is so coherent that it is difficult to imagine all those pieces fitting into any other structure. Maybe a better analogy to calling relativity "just a theory" is calling an assembled jigsaw puzzle "just one way to fit the pieces together."
Gotcha, the just-a-theory crowd says, by making that analogy you are showing that you are rigid in your thinking and unwilling to accept alternative explanations. Nonsense. Scientists are constantly trying to prove accepted theories wrong. Anyone who succeeds in disproving relativity, the Big Bang, or evolution will win a Nobel Prize and eternal fame, so we'd be happy to do so. But we know from experience that the most likely explanation for an isolated fact that seems to contradict relativity, the Big Bang, or evolution is that the fact itself was taken out of context or is not being properly interpreted, rather than that an extremely well-tested theory is wrong.
This doesn't mean that we will twist any fact to make it fit into our well-accepted theories. It does mean that surprising facts may end up extending the theory rather than replacing it. For example, Newton's theory of gravity explains a ton of observations about the motions of the planets and stars, but in a few extreme circumstances (such as very close to the Sun) it doesn't predict exactly what is observed. Einstein developed a theory of gravity (general relativity) which does correctly predict these situations. Einstein's theory is more complicated than Newton's, but in most situations the complicated parts of Einstein's theory have very little quantitative effect so we can simplify it a great deal and in those cases it turns out to be identical to....Newton's theory! This almost had to be the case, because Newton's theory accounted so well for so many observations that it would be hard to imagine that it was wrong rather than incomplete.
This example shows that a small number of facts can be critically important and that scientists do pay attention to facts which don't fit the theory. But we don't modify or overturn theories willy-nilly. When the planet Uranus didn't move exactly as Newton's theory predicted, modifications of the theory were considered but so was the possibility that some mass other than the Sun and the known planets was pulling on Uranus, and that led to the discovery of Neptune. If we rejected well-established theories at the first hint of any discrepancy with new observations, we would be giving undue weight to the new observations and too little weight to the vast range of previous observations explained by the theory. If you want to overthrow a theory because some new observation seems to contradict it, then give us a better theory which explains the new observation while still fitting the previous observations just as well as the old theory. That latter part seems to be conveniently forgotten by people who want to reject well-established theories.
A closely parallel situation is that of criminal investigators and prosecutors who present their "theory of the crime" to a jury. ("Model of the crime" would better fit my vocabulary hierarchy, but this is the word actually used.) A lot of facts may be introduced into evidence ("a car with the suspect's license plate was recorded crossing the Tappan Zee Bridge at 2:20am on August 31"), but by themselves they don't mean anything important. A good theory of the crime provides a coherent explanation of so many different facts that the jury is forced to conclude that it is true beyond a reasonable doubt. If you want to call it "just a theory" then offer us a different theory which fits the facts just as well. The defense is given sufficient time and strong motivation to offer a good alternative theory, so failure to present one is damning.
Subscribe to:
Posts (Atom)












