Category: History of Physics

  • Heaviside’s Other Equations

    Oliver Heaviside is remembered, when he is remembered at all, as the man who compressed Maxwell’s twenty equations into four. It is a reputation for tidying up: taking someone else’s theory and putting it in order.

    In 1893, in a short piece in The Electrician, he did something else entirely. He proposed that gravitation might obey equations of the same form as electromagnetism — and worked out some of what would follow if it did.

    The analogy

    The starting point is a resemblance anyone can see. Newton’s law of gravitation and Coulomb’s law of electrostatics have the same shape: an inverse-square force, proportional to the product of two quantities, one attracting masses and the other charges. Mass plays the part of charge.

    The resemblance had been noticed for a century and a half and had led nowhere in particular. What Heaviside asked was different. Electrostatics is only the static corner of electromagnetism; the full theory has a magnetic field, and disturbances that propagate. If gravity matches the corner, does it match the rest?

    Suppose it does. Then three things follow immediately, and each of them was, in 1893, a startling claim.

    There would have to be a gravitational counterpart of the magnetic field — a second field, produced not by mass at rest but by mass in motion, acting on other moving masses in a way that has no place in Newton’s theory. Heaviside had no name for it; the modern term is gravitomagnetism.

    Gravitational disturbances would have to propagate at a finite speed — and, if the analogy holds all the way, at the speed of light. In Newtonian gravity the force is instantaneous everywhere, and every attempt to give it a delay had run into trouble.

    And there would have to be a gravitational Poynting vector: a definite statement about where gravitational energy flows and at what rate. Heaviside wrote it down.

    What he did not claim

    What makes the paper interesting to read now is its restraint.

    Heaviside does not present this as a theory of gravitation. He presents it as a question about an analogy, followed by its consequences. And when he arrives at the gravitational energy flux, he admits that he does not understand what that energy actually is — the same admission Maxwell had made about the electromagnetic field, and for the same reason. The formalism produces a quantity that behaves like an energy flux. Whether something is really flowing, and if so what, the formalism does not say.

    This is worth dwelling on, because it is precisely the discipline that separates a useful analogy from a crank one. Heaviside had a construction that reproduced Newtonian gravity in the static limit and predicted new effects outside it. He could have argued that the new effects were therefore real. He argued instead that they would be real if the analogy held, and that whether it held was an open question he was not in a position to settle.

    Twenty-two years early, and on the wrong foundation

    Einstein completed General Relativity in 1915. Take the full theory, assume the gravitational field is weak and the velocities small, and linearise: what comes out is a set of equations formally very close to Maxwell’s, with a gravitoelectric field reproducing Newtonian gravity and a gravitomagnetic field generated by moving mass. This is the modern subject of gravitoelectromagnetism, and it is standard, uncontroversial physics — a limiting case of General Relativity, not an alternative to it.

    Heaviside got there in form, twenty-two years earlier, from an analogy.

    The word form is carrying weight. What Einstein’s theory says is that gravity is the curvature of spacetime, and the Maxwell-like equations emerge as an approximation valid when curvature is small. Heaviside had no spacetime, no curvature, and no principle from which his equations followed; he had a resemblance and the nerve to push it. The equations look alike. What is underneath them does not.

    Which is why the right description is not that Heaviside anticipated General Relativity. He anticipated one of its approximations, without the theory it approximates.

    The measurement

    The gravitomagnetic effect is not a formal curiosity. It is real, it is small, and it has been measured.

    A rotating mass drags the inertial frames around it — the Lense–Thirring effect, worked out in 1918. A gyroscope in orbit around the Earth should therefore precess, by a tiny amount, in the direction of the Earth’s rotation. Gravity Probe B was built to detect it: four cryogenic gyroscopes in polar orbit, launched in April 2004, collecting data from August 2004 to August 2005, with the final analysis published in 2011 after five further years of work.

    The frame-dragging drift came out at −37.2 ± 7.2 milliarcseconds per year, against a General Relativity prediction of −39.2. A milliarcsecond is about five billionths of a radian; the effect is roughly one part in a hundred and eighty of the much larger geodetic precession measured alongside it.

    It is worth being exact about what this confirms. It confirms General Relativity, whose prediction it matches. It does not confirm Heaviside’s 1893 equations, which are not the theory that predicted the number. What it establishes is that the kind of effect Heaviside reasoned his way to — a gravitational field generated by rotation, with no Newtonian counterpart — is a real feature of the world.

    What the episode is good for

    There is a version of this story that overclaims, and it is easy to write: the self-taught outsider who saw further than the establishment, whose ideas were confirmed a century later. Heaviside’s biography supports the telling — he worked outside the institutions and spent much of his life short of money — and it would be nearly true.

    Nearly true is the problem. He did not have General Relativity. His equations rest on nothing except an analogy, and analogies of that kind fail at least as often as they succeed — the same period produced mechanical models of the ether that led nowhere at all.

    The accurate version is more useful anyway. A structural resemblance between two theories is a legitimate thing to follow, and following it can put you in the right neighbourhood decades ahead of the physics that justifies being there. It does not put you in the right house. Heaviside knew the difference, which is why he wrote down what would follow if the analogy held, and stopped.

    That is a harder discipline than it sounds, and it is the reason the 1893 paper still reads well.


    Sources

    • O. Heaviside, “A Gravitational and Electromagnetic Analogy,” The Electrician, vol. 31, pp. 281–282 and 359, 1893.
    • J. Lense and H. Thirring, “Über den Einfluss der Eigenrotation der Zentralkörper auf die Bewegung der Planeten und Monde,” Phys. Z., vol. 19, pp. 156–163, 1918.
    • B. Mashhoon, “Gravitoelectromagnetism: A Brief Review,” arXiv:gr-qc/0311030.
    • C. W. F. Everitt et al., “Gravity Probe B: Final Results of a Space Experiment to Test General Relativity,” Phys. Rev. Lett., vol. 106, 221101, 2011.
  • Lorenz, Not Lorentz

    There is a condition, written down in 1867, that every student of electromagnetism meets sooner or later. It relates the two potentials of the electromagnetic field, and it is the reason those potentials can be made to propagate causally, at the speed of light, rather than adjusting themselves instantaneously across all of space.

    ∇⋅𝐀+1c2∂φ∂t=0\nabla \cdot \mathbf{A} + \frac{1}{c^2}\frac{\partial \varphi}{\partial t} = 0

    It is named after the man who published it. Almost everyone attributes it to somebody else.

    What a gauge is

    Readers who already work with gauge freedom can skip to the next section. For everyone else, the idea is worth setting out plainly, because the whole story turns on it.

    Electromagnetism can be described in two ways. One uses the fields directly — the electric field E and the magnetic field B, the things that push on a charge and that an instrument can measure. The other uses potentials: a scalar φ and a vector A, from which the fields are obtained by differentiation.

    The awkward fact about the second description is that it is not unique. Given any charge and current distribution, there is no single correct pair (φ, A). There are infinitely many, and they all describe exactly the same physical situation. Take any sufficiently smooth function χ of position and time, and transform:

    𝐀→𝐀+∇χ,φ→φ−∂χ∂t\mathbf{A} \rightarrow \mathbf{A} + \nabla\chi, \qquad \varphi \rightarrow \varphi – \frac{\partial \chi}{\partial t}

    The potentials change. The fields do not. Every measurable consequence is identical.

    The closest everyday parallel is the choice of a zero for height. To work out the potential energy of a book on a shelf, you have to decide where you are measuring from — the floor, the ground outside, sea level, the centre of the Earth. Each choice gives a different number, and none of them is the right one. But the difference in energy between the shelf and the floor comes out the same no matter which you pick, and that difference is the thing you can actually measure when the book falls. The zero is a convention. The drop is physics.

    Gauge freedom is the same situation, with a much larger space of conventions available. Choosing a gauge means picking one member of the family and agreeing to work with it. A quantity is called observable when it comes out the same regardless of that choice — and only observables can correspond to something an experiment could detect. A quantity that changes when you change gauge is an artefact of the bookkeeping, however useful it may be along the way.

    The condition at the top of this page is one such choice. It does not add physics. It fixes a convention — and, as it turns out, a particularly good one.

    Copenhagen, 1867

    Ludvig Valentin Lorenz was born in Helsingør in 1829 and became the first Danish theoretical physicist to gain an international reputation. In 1867 he published, in the Philosophical Magazine, a paper with the title On the Identity of the Vibrations of Light with Electrical Currents.

    He was working from Maxwell’s theory, which he had read closely. What he obtained were general integral solutions of the field equations in which the finite speed of light is built in explicitly: the potential at a point now depends on what the sources were doing earlier, at a time separated by exactly the light-travel delay. These are the retarded potentials.

    φ(𝐫,t)=14πε0∫ρ(𝐫′,t′)|𝐫−𝐫′|dV′𝐀(𝐫,t)=μ04π∫𝐉(𝐫′,t′)|𝐫−𝐫′|dV′\varphi(\mathbf{r},t) = \frac{1}{4\pi\varepsilon_0}\int \frac{\rho(\mathbf{r}’,t’)}{|\mathbf{r}-\mathbf{r}’|}\,dV’ \qquad\qquad \mathbf{A}(\mathbf{r},t) = \frac{\mu_0}{4\pi}\int \frac{\mathbf{J}(\mathbf{r}’,t’)}{|\mathbf{r}-\mathbf{r}’|}\,dV’
    t′=t−|𝐫−𝐫′|ct’ = t – \frac{|\mathbf{r}-\mathbf{r}’|}{c}

    The gauge condition falls out of this construction rather than being imposed on it. That is the point that gets lost when it is described as one convenient choice among several: it is the choice under which the potentials themselves respect causality, propagating outward at c like everything else in the theory.

    Maxwell, for his part, did not receive the work warmly.

    Göttingen, 1861

    There is an earlier thread, reported by the historian Edmund Whittaker in his History of the Theories of Aether and Electricity. Bernhard Riemann had discussed substantially the same condition in lectures in 1861 — six years before Lorenz, and never in print.

    This is a familiar shape in the history of science, and it cuts in an unexpected direction here. It does not diminish Lorenz’s claim; it strengthens it. Priority in science attaches to publication, not to private insight, precisely because publication is what makes an idea available to everyone else. Riemann thought of it. Lorenz put it where it could be used.

    Which makes what happened next harder to excuse.

    The other Lorentz

    Hendrik Antoon Lorentz was Dutch, born in 1853, and a physicist of the first rank in his own right: the transformations that bear his name, the electron theory of matter, the Nobel Prize in 1902 shared with Zeeman. He is one of the figures without whom twentieth-century physics does not happen.

    He also had nothing to do with the gauge condition. When Lorenz published it in 1867, Lorentz was fourteen years old.

    The attribution nonetheless drifted to him, and stuck, for two reasons that reinforce each other. The first is simply that the names sound alike and differ by one letter. The second is more interesting: the Lorenz condition happens to be Lorentz-invariant. It is compatible with special relativity, unchanged in form under a change of inertial frame — which is exactly what one would expect of something Lorentz had produced, and which is in fact a coincidence of a Danish physicist writing four decades before relativity existed.

    An error that plausible is very hard to dislodge. It appears in serious scientific literature and in reference texts, Feynman’s among them, and it has needed a small specialist literature of its own to document: van Bladel in 1991, Nevels and Shin in 2001, and the historical review by Jackson and Okun in the same year.

    It is a clean case study in how attribution actually works. A name does not win in collective memory by being correct. It wins by being familiar, and by fitting the story the reader already expects.

    Why the choice matters

    Gauge freedom means no choice is more true than another. It does not mean every choice is equally convenient, and the differences are substantial.

    The Coulomb gauge sets the divergence of A to zero. It reduces the equation for the scalar potential to an ordinary Poisson equation, which makes electrostatic problems straightforward — at the cost that φ then appears to respond instantaneously to a change in the charge distribution anywhere in the universe. Nothing observable propagates faster than light; the apparent instantaneity is confined to the potential and cancels in the fields. But you have to know that, and keep track of it.

    The Lorenz gauge does not have that feature. Both potentials satisfy wave equations, both propagate at c, and causality is manifest in the formalism rather than something one has to argue for afterwards.

    If gauges are lenses, this is the sense in which they differ: not that one shows more of the object, but that each is ground for a different kind of work. The Coulomb gauge is the right lens for a static problem. The Lorenz gauge is the right lens for anything that has to travel.

    It is worth calling it by the name of the man who published it.


    Sources

    • L. V. Lorenz, “On the Identity of the Vibrations of Light with Electrical Currents,” Phil. Mag., series 4, vol. 34, pp. 287–301, 1867.
    • E. T. Whittaker, A History of the Theories of Aether and Electricity, vol. 1. London: Nelson, 1910; rev. 1951.
    • J. Van Bladel, “Lorenz or Lorentz?,” IEEE Antennas and Propagation Magazine, vol. 33, no. 2, p. 69, 1991.
    • R. Nevels and Chang-Seok Shin, “Lorenz, Lorentz, and the Gauge,” IEEE Antennas and Propagation Magazine, vol. 43, no. 3, pp. 70–71, 2001.
    • J. D. Jackson and L. B. Okun, “Historical Roots of Gauge Invariance,” Rev. Mod. Phys., vol. 73, pp. 663–680, 2001.
  • Did Maxwell Write in Quaternions?

    On the afternoon of 16 October 1843, walking with his wife along the towpath of the Royal Canal in Dublin, William Rowan Hamilton found the thing he had been looking for over several years: a four-dimensional number system, closed under multiplication, obeying the rule i² = j² = k² = ijk = −1. He carved it into the stone of Broom Bridge on the spot. A plaque there still records the moment.

    Twenty-one years later, Maxwell read his dynamical theory of the electromagnetic field to the Royal Society. It is often said that he wrote that theory in Hamilton’s algebra, and that Oliver Heaviside later tore the quaternion structure out and replaced it with the vector calculus every student now learns. The story is appealing: a richer mathematics, discarded for convenience, with something lost in the exchange.

    It is also, in the form usually told, wrong — and the accurate version is more interesting.

    What the 1865 paper actually contains

    The paper has a well-documented history. Maxwell sent it to the Royal Society on 27 October 1864; it was read on 8 December of that year; it went to William Thomson for review before Christmas and to George Gabriel Stokes the following March; the Committee of Papers approved it on 15 June 1865 and it went to the printers the next day. It appeared as A Dynamical Theory of the Electromagnetic Field in Philosophical Transactions, volume 155, pages 459–512. The original manuscript — 84 pages in Maxwell’s hand — is still in the Society’s archives.

    Open it, and there are no quaternions.

    What there is instead is twenty equations in twenty unknowns, written out component by component in Cartesian coordinates. Faraday’s law occupies three separate lines, one each for the components Maxwell calls P, Q and R. The Ampère–Maxwell law occupies three more, for α, β and γ. Derivatives appear one axis at a time: d/dx, then d/dy, then d/dz. There is no Hamiltonian product anywhere, no scalar-plus-vector object, nothing that requires the algebra carved on Broom Bridge.

    The count itself is the giveaway. Twenty equations are needed precisely because each vector relation has to be written three times over. With a compact formalism — any compact formalism — four will do, which is exactly what Heaviside would later demonstrate.

    This does not mean Maxwell lacked the vector idea. He plainly had it. The quantities we now write as single bold letters are there in his text as triples of separate symbols: the electric field as (P, Q, R), the magnetic induction as μ(α, β, γ), the vector potential as (F, G, H). He knew these belonged together. He had no notation in working order that would let him say so in one line.

    Worth noting in passing: the twenty equations also include Ohm’s law, the force on a moving charge, and the continuity of electric charge — relations nobody today calls one of Maxwell’s equations. The set we now teach under his name was not only re-notated after his death. It was reselected.

    1870, and the Treatise

    Maxwell’s interest in quaternions was real, and it came later.

    In November 1870 he wrote a manuscript specifically on the application of quaternions to electromagnetism, and he corresponded on the subject with Peter Guthrie Tait — Hamilton’s collaborator, and the most committed advocate the algebra had in Britain. In the Treatise on Electricity and Magnetism of 1873, a more condensed quaternion notation appears for the general field equations, in the second volume.

    So the honest sequence is: no quaternions in 1865, a manuscript on the question in 1870, condensed quaternion notation in the 1873 Treatise. The mythology compresses this into a single origin that never existed.

    Even in the Treatise, the use is limited. On the analysis of the historian André Waser, Maxwell employed quaternions as an expository and symbolic language rather than as an operational tool — a way of displaying the structure of a result, not a machinery he calculated in daily. Whether that judgement holds under a page-by-page reading of the relevant chapter is a question worth putting to the original text rather than taking on authority, and one I intend to return to.

    The most defensible summary is this: Maxwell was genuinely interested in quaternions, adopted them partially, and never brought their systematic use in electromagnetism to completion. He introduced the tool into the subject. He did not build with it.

    Heaviside, Gibbs, and the four equations

    Maxwell died in 1879, at forty-eight.

    The reformulation came afterwards. Through the 1880s, Oliver Heaviside — self-taught, working outside the institutions, and in financial difficulty for much of his life — developed a vector calculus and rewrote the twenty equations into the four with divergence and curl that are now universal. Josiah Willard Gibbs arrived at substantially the same system independently at Yale.

    The chronology matters for the story. There was no confrontation between Maxwell and Heaviside over notation, and Maxwell did not live to see his own formulation set aside. Whatever was decided about quaternions in electromagnetism was decided by other people, in a debate he was not present for.

    Why the vector form won is not mysterious. It was fitted to the problems that mattered at the moment it appeared. The rising technology of the period was radio, and radio is dominated by transverse waves in free space, where the vector calculus is exactly adequate and quaternions are overhead. Heaviside’s system was leaner for the work in front of it, and it worked — unambiguously, for well over a century, underwriting almost everything electrical that followed.

    The war of the vectors

    The quaternionists did not concede quietly. Through the 1890s, in the pages of Nature and elsewhere, Tait fought a long rearguard action against Gibbs and Heaviside, defending Hamilton’s algebra as the proper language of physics against what he regarded as a mutilation of it.

    He lost, comprehensively. By the turn of the century the vector notation was standard in physics teaching, and quaternions had been pushed to the margins of the curriculum, where they remained until computer graphics and spacecraft attitude control found a use for them in representing rotations — a role that has nothing to do with the one Tait was arguing for.

    The point worth holding onto is who was arguing. The defence of quaternions in electromagnetism was mounted by Tait, not by Maxwell, and it was mounted after Maxwell was dead.

    Abandoned, not refuted

    Nothing in this history involves anyone showing that the quaternion approach to electromagnetism was wrong. No calculation failed, no prediction came out false, no inconsistency was exposed. The programme was not defeated. It was set down, by a man who had other things to finish, and never picked up again — because within a few years something more convenient had arrived for the problems then at hand.

    That is a weaker claim than the mythological version, and a more durable one. It also leaves a question standing that the mythological version answers too quickly. If a formalism is put aside because it is heavier than the problems of the day require, the natural thing to ask is what happens at problems it was never tested against — and whether the reduction that made the vector calculus so efficient for transverse waves in free space discarded anything along with the notation.

    That question does not answer itself, and it is not answered here. But it is a real question, and it is worth keeping the history straight in order to ask it properly.


    Sources

    • J. C. Maxwell, “A Dynamical Theory of the Electromagnetic Field,” Phil. Trans. R. Soc., vol. 155, pp. 459–512, 1865. Read 8 December 1864.
    • Manuscript, Royal Society Archives, PT/72/7 — received 27 October 1864.
    • J. C. Maxwell, A Treatise on Electricity and Magnetism, vol. 2. Oxford: Clarendon Press, 1873, Ch. IX.
    • A. Waser, “On the Notation of Maxwell’s Field Equations,” AW-Verlag, 2000.
    • O. Heaviside, Electromagnetic Theory, 3 vols. London: The Electrician, 1893–1912.
    • W. R. Hamilton, “On quaternions,” Proc. R. Irish Acad., vol. 3, pp. 1–16, 1847.
    • M. Longair, “A commentary on Maxwell (1865),” Phil. Trans. R. Soc. A, vol. 373, 20140473, 2015.