Saturday, March 6, 2021

Electrical Current

This is a question on many student’s minds: what are the electrons actually doing inside a current-carrying wire? 

The answer might seem straightforward at first glance. An electrical current is a flow of charge moving through an electrical conductor or through space. From electronics-notes.com, we have a practical definition of electrical current: it is “the rate of change of flow past a given point in an electric circuit.” From physicsclassroom.com (an excellent teaching site), the definition alludes to a bit of scientific history: An electric current is “by convention the direction in which a positive charge would move.” Electrons “move through the wires in the opposite direction.” Confusing for many students, conventional current in a wire moves from the positive terminal of a battery to the negative terminal. Electrons, being negative charge carriers, actually move in the opposite direction. This might be where many of us stop in our quest to understand current, which is unfortunate because this is a fascinating exploration.

 

The silver lining of this conventional terminology is that it nudges us toward the history, which shows an incredible advancement in solid-state physics. How was electrical current discovered? We probably all know about Benjamin Franklin’s famous kite experiment in 1752. We might not know that he didn’t actually discover electricity with this experiment, nor was he the first to discover that lightning was actually a type of electricity. Electrical forces had been known for hundreds of years by his time. I think, however, we can fairly credit him for contributing to our confusion about current. He studied static electricity  by producing a static charge on the surface of glass, amber and other materials by rubbing them with fur or a dry cloth. This resulted in an exchange of electrons from one material to another. At that time, electrical current was called “electrical fluid,” and as such, Franklin guessed that some materials (such as glass that’s rubbed) contained more of this fluid than others. To his thinking, these charged objects contained excess, or positive, electricity, while others contained a deficiency of the fluid, or negative electricity. Electric batteries were developed soon afterward and it seemed natural to assign the direction of electrical flow from positive to negative (excess to deficient). It was only when electrons, the subatomic particles responsible for static charge, were discovered about one hundred years later, that scientists realized these particles move in the opposite direction. It is an excess of electrons that produces a negative charge, so the flow from excess to deficient must actually be from a negative terminal to a positive terminal. 

 

This “conventional  current” (positive to negative) had staying power. The conventional terminal designation is still used worldwide. It’s not a problem to work with as long as it is consistent, but it can present a problem when we try to understand what is actually happening inside the conductor.

 

What then is going on inside an electrical current-carrying wire, for example? Consider an everyday power cord on a vacuum cleaner. If we could zoom into a cross-section of that cord, we would see wires made of copper through which electrons travel easily and with little resistance, surrounded by material that resists current flow and provides good electrical insulation. We can imagine electrons flowing from the electrical outlet in the wall, through the cord, and into the appliance. But where do these flowing electrons end up? Do they get used up in the process of doing work somehow? When the power is shut off, are we left with an alarming reservoir of electrons somewhere inside the vacuum motor? Where do the electrons in the wall socket originate from? These great questions start us on our journey. 

 

When we think of electrical current as a physical flow of negatively charged particles through a material or through space, we might think of something analogous to water molecules flowing in a stream, and when we do, a number of pressing questions come to mind. Like the out-dated convention of positive to negative terminal current flow, the terms “current” and “flow” themselves lead us away from clear modern evidence that electrical current is a not a physical flow of particles at all. Many physics classrooms begin their discussion of electrical current with a water analogy and it is a good place to go to get a feel for how simple electrical circuits work. But this analogy, while a good start, proves misleading as we deepen our understanding. And there is much more to this fascinating story than this. To learn it we must upgrade our understanding of what electrons are doing at the subatomic level inside a conductor.

 

What Is an Electron?

 

All materials are made of atoms. All atoms consist of a nucleus surrounded by electrons. The nucleus is composed of neutrally charged particles straightforwardly called neutrons and positively charged protons. It therefore has a positive charge. These particles are bound tightly together by a fundamental force called the strong force (strong force). This force is indeed strong. It easily overcomes the repulsive forces between the positively charged protons, but it only acts over an extremely short distance, at the scale of the nucleus itself, and from there, its influence drops off dramatically to zero. Negatively charged particles called electrons surround the nucleus. They are attracted to the nucleus through the attraction of opposite electrical charges. An electrically neutral atom contains equal numbers of electrons as protons. 

 

We might imagine electrons moving around the nucleus in planet-like circular orbits, except that here, the atomic force is electrostatic rather than gravitational. This is the familiar Bohr model (below) introduced by Niels Bohr and Ernest Rutherford in 1913. This model of the hydrogen atom shows three possible energy shells for the electron. At n=1, the electron is at its lowest energy (ground) state. If the atom is in an excited state, the electron will be in a higher energy shell. It will emit a photon of light as it returns to a lower energy state.


The idea that electrons orbit the nucleus in specific stable orbits came from necessity. The researchers knew that atomic electrons can release energy in the form of electromagnetic radiation (light) but they also knew that if an electron loses energy, it should quickly spiral into the nucleus, and in the process it would emit increasingly high frequency radiation (called the ultraviolet catastrophe) No atom would be stable for more than a few trillionths of a second. They also knew, from experiments a few decades prior, that atoms emit light only at specific frequencies. An excited (high energy) hydrogen atom will emit only purple, blue, aqua or red (the Balmer series emission spectrum, shown below), depending on the energy level of its excited-state electron, as it returns to its rest state. 


In an excited pure hydrogen gas, you will see the whole spectrum but hydrogen will never emit green or yellow, for example. These scientists figured out that electrons in atoms must occupy discrete energy levels. They orbit at specific stable distances from the nucleus. The farther an electron is form the nucleus, the higher its energy level. An electron can move up or down only by jumping to/from specific energy levels. Energies are therefore quantized. They come in quanta or packets.  This model is useful for predicting the spectral phenomena of simple atoms with few electrons, like hydrogen, but it cannot explain the spectra of large complex atoms. Nor can it explain the different intensities of spectral lines for any given atom. It is also still used as a simplified model for chemical bonding between atoms, in which atoms share one or more electrons located at their outermost energy levels.

 

The modern model of the atom, the quantum physical model, describes the positions of atomic electrons not as being precisely located within energy shells but as clouds of probability. The hydrogen atom is shown below right. The shapes of the electron orbitals are shown in yellow and blue. Denser regions of colour indicate a higher probability of the electron's location. The energy shells (1, 2 and 3) are arranged in increasing energy from top to bottom. The orbital shapes are s, p and d, shown from left to right.


This model incorporates the fact that we no longer understand electrons to be just point charges, or tiny charge-carrying billiard balls. Under certain circumstances they do act as such. For example, an electron has a measurable momentum and it can take part in elastic collisions. But electrons also act as waves and that wave nature also shows up under certain circumstances, such the creation of wavelike interference patterns. We now think of the electron as both particle and wave, which is not easy to grasp. These qualities seem to be mutually exclusive. Electrons in conductors, however, do display both wave and particle behaviours. To predict such behaviours, we must go quantum and understand electrons as wave functions that exhibit both particle AND wave behaviour. An electron’s position and momentum (or velocity) are now defined as probabilities. Where and how fast an electron is going are assigned probability amplitudes, rather than specific values. Furthermore, thanks to Werner Heisenberg’s uncertainty principle of quantum mechanics, these are complementary variables, which means we can only know one value at the expense of another (this part was actually formulated by Bohr). If we want to know precisely where an electron is, we cannot know its momentum at that same moment, and vice versa. Likewise, the electron’s wave and particle properties are also complementary. A single electron cannot simultaneously exhibit both its full wave-like and particle-like nature, with the exception of the famously fascinating double-slit experiment, in which electrons show some of both behaviours at the same, which I encourage you to look up.

 

This is a difficult conceptual step to make but we must because the easy-to-visualize Bohr model, as useful as it is, leaves something missing. By using Schrodinger’s equation to mathematically describe atomic electron behaviour as a wave function, physicists can predict many of the spectral phenomena that the Bohr model cannot. It’s not easy to do in practise either; the calculations are very complex.

 

What is an Electrical Conductor?

 

The atoms that make up a good electrical conductor have an atomic structure that allows the outermost electrons, called valence electrons, to be loosely bound to the nucleus. Valence electrons are outermost energy shell electrons. These are the electrons that take part in chemical bonding. Now we can describe chemical bonding more precisely, in quantum mechanical terms, as two or more atomic orbitals combining to form a molecular orbital. An atomic orbital is a region of space around the nucleus where an electron is most likely to be. It can be a simple spherical shape or a more complex shape depending on its energy level. The orbital shapes come from solving the Schrodinger equation for electrons bound to their atom through the electric field created by the nucleus. The orbital is part of the electron wavefunction that describes the electron’s location boundary and its wavelike behaviour.

 

Molecular orbitals, on the other hand, come in three different kinds: 1) bonding orbital, which has a lower potential energy than the atomic orbitals it is formed from, 2) antibonding orbital, which has a higher energy than the atomic orbitals, so it opposes chemical bonding and 3) nonbonding orbital, which has the same energy so it has no effect on chemical bonding either way. A chemical bond is a constructive, in-phase, interaction between two valence electrons of two atoms. An antibonding orbital is an interaction between atomic orbitals that is destructive and out of phase. The wavefunction of an antibonding orbital is zero between the two atoms. This means there are no solutions to the Schrodinger equation for this orbital and therefore there is no probability of an electron being available to bond. 

 

Metals tend to be good conductors because their valence electrons are loosely bound to their nuclei. They are delocalized electrons, which means that they are not associated with a particular atom or chemical bond. These molecular orbital electrons extend outward over many atoms. Although the electrons are delocalized, the metal atoms themselves are bound tightly together through metallic bonding, by electrostatic attractions between the positive nuclei and the “sea” of electrons in which they are embedded. The atoms are held tightly together, which means metals tend to have high melting and boiling points. The nuclei in metals act like positive ions in this arrangement, which means that metallic bonding is similar in this sense to ionic bonding. The resulting metal structure is a tight three-dimensional lattice arrangement, similar to the atomic structure of ionic crystals such as sodium chloride (table salt). In contrast to ionic compounds, valence electrons in the metallic lattice form molecular orbitals that extend across the entire metal. The bonding electrons themselves do not orbit the entire metal but their influence extends across the metal. Valence electrons in metals act more like a collective than they do in conventional chemical bonds. 

 

All of the valence electrons in the metal participate in molecular bonding. As vast as this collection of delocalized electrons is, the number of possible delocalized electron energy states is far greater. All metal atoms contain few electrons in their valence energy orbitals. Transition metal atoms, for example, can hold up 18 electrons in the outermost energy shell (which consists of 5 d orbitals, one s orbital and three p orbitals), but they are barely filled. If this is confusing, we will be exploring this in more detail later. The point here is there are many more possible energy states than there are electrons available to fill them. These empty available states, which are all similar in terms of energy, will become important when we look at band theory later on in this article.

 

What Happens When an Electric Potential is Applied?

 

When a metal with high conductivity is placed in an electric field, all the valence electrons tend to move against the direction of the field. An electric field is a vector force field that surrounds an electrical charge and exerts a force on other charges nearby. The electrons themselves each generate an electric field, as do the positively charge nuclei in the metal. When electrons within the metal move, they also generate a magnetic field. We can begin to see electric current as a complex interplay between multiple local electric and magnetic fields.

 

An analogy of a waterfall is often used to describe the movement of charged particles, such as electrons, as moving along or down an electric potential. An electric potential might at first sound the same as an electric field but the electric potential expresses the effect of an electric field at a particular location in the field. We could place a test charge within an electric field and measure its potential energy as a result of being in that field. (A positive test charge is usually used and this is why electrons move against the direction of an electric field.) The potential energy of the electric field will differ, depending on the location within that field. A negative test charge, for example, will have high potential energy close to a negative source charge and lower energy further away from it. In other words, the electron tends to move from where it would have high potential energy toward lower potential energy. Inside a copper wire attached to a battery, for example, the electron is pushed away from the high-energy negative terminal (labelled positive!) and toward the region of lowest electric potential energy, the positive terminal (labeled negative!). This difference in electric potential energy is called voltage. Voltage is defined as the amount of work required per unit of charge to move a test charge between two points. 1 volt = 1 joule of work done per 1 coulomb of charge. For example, if a 12-volt battery is used in a wire circuit to power a light bulb, every coulomb of charge in the circuit gains 12 joules of potential energy it moves through the battery. Every coulomb in turn loses 12 joules of energy into the environment as light (and some heat) as it powers the light bulb.

 

The waterfall analogy, as mentioned earlier, can be misleading. Within the wire, the electrical current does not depend on electrons wriggling free from copper atoms and flowing down the wire, like water molecules would flow down a waterfall. The description of a “delocalized sea of electrons” can naturally lead to such an assumption about current. Instead, each valence electron is held loosely enough to its nucleus to “nudge” a neighbouring electron in a neighbouring atom and so on. A better analogy for electrical current might be that of fans in a football stadium row standing up one after another to do the wave. The people stay in place but the wave moves down the row. If we want to stick with water, we could say that the current is analagous to the wave travelling across a sea. Electric current can be defined as the rate at which charge (not electrons) flows past a point in a circuit. 1 ampere of current = 1 coulomb of charge/second.

 

It’s easy to imagine a flow of electrons under pressure (an analogy for voltage) gliding along inside a wire like water flowing through a pipe, but current is not the physical flow of electrons, but rather the flow of the energy of their movement. We can put this idea more scientifically as a momentum transfer. Put yet another way, it is a transfer of kinetic energy from one electron to another and another and so on. We can think of nudged electrons transferring energy from one to the next like a billiard ball hitting an adjacent billiard ball along a line of billiard balls. This analogy alludes to the particle nature of the electrons.

 

Electrical Conduction: The Drude Model

 

The “billiard ball” description of electrical conduction is called the Drude model, proposed in 1900. This model was a bit before it’s time because it predated even the (Rutherford) atomic model by a few years. The “sea of electrons” embedded within a positive matrix that Rutherford envisioned (also called the raisin bun or plum pudding model) just happens to align pretty well with the Bohr valence electron model for metals, where a sea of valence electrons surrounds each atom. The Drude model is essentially  a description of how kinetic energy is transferred within a conductor by treating electrons like tiny billiard balls. For those science history buffs like me, we can follow how the Drude model evolved into our modern model. We can read into it how our understanding of electron behaviour evolved. The Drude model was not a modern quantum mechanical model of what is happening, but it could predict some electron conduction behaviour by understanding conduction as a momentum transfer between electrons exhibiting particle-like behaviour. In fact, Paul Drude used Maxwell-Boltzmann statistics to derive his model. These statistics describe an average distribution of non-interacting particles at thermal equilibrium. This is a classical description of how an ideal gas behaves, and it was the model available to him at the time. In this view, each atom in a conducting metal such as copper contributes valence electrons to a sea of non-localized, non-interacting electrons. It turns out he was accidentally right. In most cases, at least in the case of copper, we can effectively neglect the interactive forces between the electrons because they are shielded from those forces.  All the relatively stationary (much more massive and more highly charged) atomic nuclei in the copper present an overwhelming influence on them. It is this shielding effect) that allows us to model metal electrons fairly accurately as an ideal gas, a cloud of particles that do not interact with each other. However, the particles, in this case electrons, are far more concentrated together than the atoms in any gas would be. This means that this model doesn’t predict all conduction behaviour. Electrical conduction exhibits behaviours that are too complex to be described through classical ideal gas theory. 

 

Maxwell-Boltzmann statistics were replaced around 1926 by Fermi-Dirac statistics. Rather than focusing on the non-interaction between particles, these statistics describe a distribution of particles over a range of energy states in a group of identical particles that all obey the Pauli exclusion principle. This effectively upgraded the rule for interaction. Rather than saying the electrons don’t interact, we now say they cannot occupy the same quantum state. We now have to treat electrons as quantum particles. Everything about a quantum particle is described by four quantum numbers (spin, magnetic, azimuthal and principle). Electrons belong to a special division of quantum particles called fermions. Fermions can share any three quantum numbers but they can’t overlap in the same location at the same time. A boson such as a photon can, and it is governed by a different set of statistics. We are introducing the Pauli exclusion principle to the gas-like behaviour of valence electrons inside a metal. The Pauli exclusion principle is a critically important rule that electrons in atoms (and all fermions) must obey. It is a quantum mechanical principle. The implications are quite profound. It means that the potential energy of the gas-like electrons must now be taken into account, and by doing so, it limits the numbers of electrons that occupy each orbital in an atom. This rule forbids valence electrons from moving down to occupy already filled lower energy states in the atom. This means there is a lower limit on the  potential energy of the atom’s (lowest energy) ground state. By incorporating quantum mechanical effects, this model significantly improved the predictions we can make about conducting electron behaviour. 

 

Band Theory of Metals

 

The Drude model, as good as it is, is still missing a component we need to take into account with conduction in metals.  We need to consider the effects of how molecular bonding relates to the conductivity of the metal. We’ve been modelling the conduction electrons as particles that have limited occupied energy states but otherwise act like strangers to each other. 

 

Molecular orbitals in metals are much larger than atomic orbitals, which are confined within the atom. In fact, molecular orbitals extend and overlap each other throughout the whole material. Bearing this in mind, the word “orbital” here can seem misleading. There is no orbiting implied here. Better said, the set of molecular orbitals that are created in the metal are a set of interactions that are generated by the valence orbitals of the interacting atoms. This molecular interaction, rather than the electrons themselves, extends throughout the material. 

 

Metallic bonding means that a very high number of atomic valence orbitals interact simultaneously within a metal. This is a critical key to explaining why metals can be such good electrical conductors. Within a free atom, an atom that is not chemically bonded to any other atom, the energy that an electron can possess must fall into one of several possible discrete energy levels. Within a conductive metal, on the other hand, due to the overlapping of a huge number of molecular orbitals, an enormous number of new possible energy states opens up. The available energies of the valence electrons are no longer confined to discrete energies but instead now to a band of available energies that can vary in width. This molecular orbital approach is called band theory. It’s a very useful way to visualize how conduction works, and to understand why some materials are better conductors than others. 

 

Why are the valence energies spread out into a band? First, we must distinguish between an orbital and a shell. Inside any atom, valence electrons are confined to the outermost (highest energy or bonding) valence energy shell. An electron shell is all about the energy. An atomic orbital is about where an electron is most likely to be found. To distinguish these two concepts, we can imagine that we are adding electrons to an atom. First we fill up a sphere-shaped 1s-orbital, which can hold two (opposite-spin) electrons. Then we fill up the next highest energy 2s spherical orbital with two electrons Then we start to fill three 2p orbitals (three dumbbells shapes in three-dimensions) with 2 electrons each, six electrons total, and so on. Every atom of every element possesses the same number of potential orbitals and shells, but they differ in the number of electrons inhabiting them and the orbitals themselves can differ in size. The 1s electrons belong to the lowest energy K shell. The 2s and 2p electrons belong to the next higher energy M shell. The M energy shell, which is just a circle in a Bohr diagram, can now described in three-dimensional detail, as a double-lobbed and spherical orbital. The three 2p orbitals (which extend along the x, y and z axis in three-dimensional space) and 2s orbital hold a total of eight electrons in the M energy shell. A simplified diagram is shown below left.


The same energy level can contain multiple orbitals. This is Hund’s rule

 

Let’s look at the electron configuration of copper as an example. A copper atom, with 29 electrons in total, has an unexpected electron configuration that hints at why it is so conductive. If we simply fill up orbitals in order (1s, 2s, 2p, 3s, 3p, 4s, 3d; a 4-lobed d orbital can contain up to 10 electrons) we will have 1s22s22p63s23p64s23d9. This is actually wrong, based on experimental evidence. For copper and other transition metals, we must write the 3d orbital before the 4s orbital even though 3d is considered to be higher energy than 4s. This is an oddity of the transition elements. In these elements, the 4s orbital behaves like an outermost highest energy orbital. This short 3-minute youtube video explains why:



Following the transition element rule, we should have 1s22s22p63s23p63d94s2, thinking (correctly) that 4s should still fill up before 3d starts to fill. This is still wrong. According to experimental evidence the correct configuration is 1s22s22p63s23p63d104s1. In these metals, the 3d orbital is just slightly larger than the 4s orbital, but it is still higher energy than 4s the same as it is in other atoms. However, it needs only one more electron to be filled (with 10 electrons total). This filled d orbital state is a more stable lower energy configuration, so it pulls up an electron from the slightly lower energy 4s orbital to achieve it. That electron, moving up from 4s to 3d, increases in energy but the increase is very small. That small cost pays off by significantly lowering the atom’s total potential energy. 

 

At first glance you might conclude that copper has one valence electron, just one electron available to delocalize and form extensive molecular orbitals that take part in electrical conduction. You cannot read it from the configuration, but the 3d and 4s orbitals are very similar in energy. That’s why the electron swap is possible. This means that in practice, copper has eleven electrons available to contribute to the valence energy shell, and therefore to molecular orbitals and to electrical conduction.

 

If we go back to our copper wire example, these valence electrons will fall into a range of many possible energy states due to the overlapping of a huge number of molecular orbitals. This range is called the valence band. In conductive metals, there are many empty valence vacancies with energies just above the filled orbitals that electrons, can move into and out. This range of energies just above the valence band is called the conduction band. In fact, in metals, these two bands of energies overlap, so that some electrons are always present in the conduction band.

 

Another energy level that is very useful to understand is the Fermi level. This can be defined as the maximum available energy level for an electron in a material that is at absolute zero. Absolute zero is the coldest possible temperature of a material. It is the lowest possible potential energy state. In a conductor at absolute zero, the valence electrons are all packed into lowest available valence orbitals. Thanks to the Pauli exclusion principle, the electrons must retain a range of base level energies. We can call this range a “Fermi sea” of electron energy states. The Fermi level is like the surface of this sea, where no electron has enough energy to rise above the surface, an analogy I lifted from the Hyperphysics link above. I think this analogy can lead to some confusion. We might assume that the Fermi level is simply the top of the valence band, which is incorrect. From Wikipedia, we can understand the Fermi level more deeply as “a hypothetical energy level of an electron, such that at thermodynamic equilibrium this energy level would have a 50% probability of being occupied at any given time.” It’s critical to note that the Fermi level does not correspond to any actual electron energy level. Instead it is an energy state that lies between the valence and conduction band energies, or where they overlap in the case of metals. In metals, there are an equal number of occupied and unoccupied energy states at the Fermi level. There are many electrons and there are many free states, and therefore, conductivity is maximized. In contrast, in a material where all the energy states are occupied at the Fermi level, there is nowhere for electrons to move to, so it cannot conduct electricity. 

 

At room temperature some copper electrons are already in the conduction band. We might guess that the conductivity of a copper wire would simply continue to increase as the temperature increases above room temperature, but the relationship is more complex. As the temperature goes up, as electrons in the copper atoms gain thermal energy, they also gain non-directional kinetic energy and therefore lattice vibrations in the copper grow. This kind of electron motion is random and “jittery.” Electrical conduction depends on the motion of electrons in one direction. As the temperature increases, electrical conduction becomes increasingly hindered by internal collisions between delocalized electrons, and conductivity therefore decreases. If we cool copper down toward absolute zero, we would expect fewer and fewer electrons with enough energy to conduct electricity as they fall below the Fermi level. At the same time, thermal vibrations within the metal lattice calm down so conducting electrons are less likely to be scattering by other electrons. If it were possible to reach absolute zero, we would expect no electrons at all above the Fermi level and there would be no conduction possible. It seems to me we could never apply an electric potential to test for current without adding the energy of an electric field that would excite some electrons into conduction. 

 

Electric Insulators

 

Every element and material has a valence band and a conduction band. In terms of electrical conductivity, materials fall into one of three categories: conductor, insulator and semiconductor. Even the best insulator is not a perfect insulator. It too has a conduction band that, under the right conditions, can be populated by a few mobile electrons. There is a significant difference between insulators, semiconductors and conductors in terms of how far apart the energies of the valence and conduction bands are. In the diagram below the energy is increasing from bottom to top. 



We can describe the valence electrons in every material as having two possible ranges of energy levels, in other words, two energy bands: a valence band and a conduction band. Between these bands is a gap, a range of energies that no valence electrons can occupy. Its width varies greatly between materials. Each element or material therefore has a unique band structure. In all insulators, there is a large range of energy above the valence band where no electron energy orbitals are available. They have large band gaps. In an insulator, the organization of the atoms does not allow for free electrons. All of the electrons are tightly bound to their nuclei (there is no shielding effect) and they form tight localized bonds between atoms. A great deal of energy must applied to an insulator before its valence electrons have enough energy to jump the band gap into the conduction band. Once enough energy is applied, however, even an insulator will start to conduct electricity. At ground state, valence electrons are tightly bound to the atoms in the insulator material. In a sufficiently high energy environment, they become excited and delocalized into conduction electrons. 

 

I am being careful to write “energy” rather than “temperature” here because a strong electric field is much better than heat for tearing valence electrons away from atoms and turning them into conduction electrons. Heat is a randomly directed influence that generates random jittery electron movements in the material. 

 

To see what happens when an insulator becomes a conductor, consider a resistor hooked up to a circuit. Let’s apply a voltage much higher than it is rated for. The electric field created by the applied voltage will eventually overcome the material’s resistance. We call this resisting property its dielectric strength. When its dielectric strength is overcome, the material breaks down into a conductor. In a solid, the breakdown is physical and permanent. Our little resistor is burned and ruined. We can also observe this breakdown as an electrostatic discharge, a familiar example being lightning. A sudden giant spark is emitted as air, normally a strong insulator, breaks down through ionization. Electrons stripped from the gas atoms become conduction electrons when a sufficient electric potential difference between one area and another builds up during a storm. 

 

Semiconductors

 

Semiconductors are widely used in electronic circuits. A semiconductor is any material that conducts current but only partially. Its conductivity falls between an insulator and a conductor. Most are made of crystals, so they have lattice structures similar to metals. In semiconductors, the band gap is smaller than in insulators, so a smaller amount of energy (a small amount of heat in many cases) is required to bridge the band gap into the conduction band. Some semiconductors can be chemically altered to further enhance their conductivity, through a process called doping. In all materials, the Fermi energy level is somewhere in the middle of the band gap and if there is no band gap, as in metals, it is near the top of the valence band.



In semiconductors, doping can shift the Fermi level either much closer to the conduction band (N type) or much closer to the valence band (P-type). Doping is done by adding a few foreign atoms into the lattice structure of the material. These impurities add extra available energy levels. In N-type semiconductors (shown below left), energy levels are added near the top of the band gap (along with additional free valence electrons that contribute to conduction). 


These additional energy levels just below the conduction band, mean that the electrons occupying them can be easily excited into the conduction band. In P-type conductors (shown below right), extra empty energy levels are created as mobile energy “holes.” 


These are holes that electrons would normally occupy in the top of the valence band of the material. Electrons can jump in and out of them. N-semiconductors are much more conductive than P-type semiconductors because there are many more electrons available in the conduction band in the N-type than there are holes in the valence band in the P-type. 

 

Conductors

 

In conductors, there is no energy gap at all. The valence band overlaps the conduction band. At room temperature, some electrons are energetic enough to be conduction electrons. When an electric potential is applied, current is generated in the material. When we want to know how good an electrical conductor is, we can ask how many electrons are available in the conduction band. Copper, as we discovered, has eleven electrons per atom that populate the valence band. These electrons are delocalized in the lattice as molecular orbitals. That’s a large number of electrons available to generate a current when an electric potential is applied. 

 

If all metals have overlapping valence and conduction bands, why are some better conductors than others? The answer can be quite complicated but we can get some idea by comparing copper with iron, which is much less conductive. Copper has a conductivity of 5.96 x 107 S/m (Siemens per metre) at 20°C. Iron’s is 1.00 x 107 S/m at 20°C, about six times lower. Copper displays an almost perfect Fermi surface which means it acts like a hypothetical free electron sphere would act. Its valence electrons are all lined up close to the Fermi level, the line between occupied and unoccupied energy shells. They move easily in any direction. Iron has two main differences. Iron atoms have a strong magnetic moment; they act like tiny magnets. This splits the band structure into two parts, based on the two magnetic spin states. This splitting separates the electrons, reducing the ways they can move. The second difference is the Fermi surface isn’t smooth for iron. A Fermi surface is an abstract geometric representation of all occupied versus unoccupied electron energy states in a metal at absolute zero. The shape is derived from the symmetry and periodicity of the metal lattice. In iron, it is broken up into many disconnected pockets, rather than a smooth electron sphere. Electrons have to jump from pocket to pocket in order to move. Even though iron has free (delocalized) electrons in molecular orbitals just as copper does, these two factors make it much harder for the electrons to move and generate current.

 

To sum up, when wondering about the conductivity of a material, a crucial question to answer is how close the conduction band is to the valence band. In a conductive material the two bands are very close and in metals they overlap. In a semiconductor at room temperature, there is a small gap in energy between the valence band and the conduction band with the Fermi level within that gap. Doping adjusts the Fermi level and may provide additional conducting electrons. In an insulator at room temperature, the energy difference between valence band and conduction band is so large that at room temperature no electrons in the valence band can absorb enough energy to populate the conduction band. If the energy of an insulator is increased (usually be applying very high voltage), some of the conductor’s valence electrons may absorb enough energy to cross the Fermi level and jump the gap into the conduction energy band. The dielectric strength of the material is overcome and the material breaks down.

 

There is a finer difference between conductors and semiconductors, in terms of the density of energy states crossing  the band energy gap. In a semiconductor, the conduction band is above the Fermi level, so as its energy goes up, the conduction band begins to get populated by electrons as they cross the gap, starting from zero to one electron to two and so on. Conductivity gradually increases starting from zero. In a good conductor, the Fermi level is in the conduction band because it overlaps with the valence band at room temperature. As the energy rises, the number of conduction electrons starts with an already populated conduction band and increases from there. This is additional reason why metal conductors conduct current so readily.

 

In addition to their conductivity, band theory also explains many physical properties of metals. 

Metals conduct heat better than other materials because their delocalized electrons can easily transfer thermal energy from one region to another. Generally, the better a metal is at conducting electricity, the better it is at conducting heat. Metals at room temperature feel cool to the touch because the electrons in the metal readily absorb the thermal energy in your warmer fingertips. This energy absorption represents a very small jump in energy to slightly higher available energy levels. For the same reason, metals under a hot sun tend to feel hotter than other objects nearby. The electrons easily transfer thermal energy to relatively cooler objects such as your fingertips. Metals tend to have a small specific heat capacity, which is a measure of how much heat must be added to a material to raise its temperature. In a similar way, metals are lustrous because numerous similar energy levels available to the valence electrons means they can absorb the energy of various wavelengths of visible light (they absorb various colours simultaneously). When electrons inevitably decay back to lower rest state energy levels, they emit light across the same wide range of wavelengths. As light is continuously absorbed and re-emitted from the surface of the metal, we see it as lustrous and shiny.

 

Three Different Velocities 

 

Now that we know how electrical conduction starts, we can focus on what’s happening when the current flows. Here we can further our quantum mechanical understanding of how the electrons in a conductor behave. A metal, as we know, is organized into a three-dimensional lattice of atoms. Each atom has an outer energy shell of delocalized valence electrons. These electrons are somewhat free of the attractive influence of their nuclei and they interact with each other, forming molecular bonds between atoms within the lattice. This “sea of electrons” is free to move and conduct an electrical current through the metal. For example, when a voltage is applied to a circuit, electrons in a copper wire drift from the negatively charged terminal toward the positive charge. The electrons themselves drift very slowly through the metal, on the order of a just few metres per hour. If we looked inside the  copper wire with the circuit switch turned off, we would see electrons continuously making microscopically short trips in random directions, changing direction as they strike other electrons and bounce off them. The net velocity of the these electrons is zero when no voltage is applied. When a voltage is applied, there is a net flow, a drift, of electrons in the opposite direction (negative to positive) to the electrical field (positive to negative) that is superimposed over the random drift motion. 

 

Wikipedia supplies an interesting mathematical example of electron drift velocity, worked out for a 2-mm diameter copper wire. The drift velocity (which is proportional to the current; here 1 ampere) works out to be 23 um/s. That’s micrometres. For a typical household 60 Hz alternating current, the electrons in the wire drift less than 0.2 um per half cycle. This means that the electrons flowing across the contact point in a switch, for example, never even leave the switch. This doesn’t mean that the electrons are individually lumbering along within the wire. Individual electrons in any material at room temperature have a tremendous amount of kinetic energy, called Fermi energy. The Fermi velocity of electrons is about 1570 km/s! The drift, the net movement of electrons in one direction under an electric potential, however, is extremely slow. The speed of electricity (the speed at which an electrical signal or energy travels) through a copper wire is yet again different. The energy that runs the motor in a vacuum travels as an electromagnetic wave from the socket, through the cord, and to the motor. The wave’s propagation speed is close to but not quite the speed of light. This is why the vacuum starts up immediately once you flip the switch. And this is where our simpler billiard ball explanation of current propagation falls down. If electrons were like tiny balls striking one another down a line, like football fans doing the wave in a stand, delays caused by the inertia of each electron as it is put into directed motion would build up quickly and significantly and slow the current to a stop. The electric signal instead travels as propagating synchronized oscillations  of electric and magnetic fields generated by the oscillations of electrons in conducting atoms.

 

The wire guides, rather than physically carries the wave of electromagnetic energy. This wave of energy, in turn, generates an electromagnetic field that propagates through space, responding to the (just preceding) energy flow. The time required for the field to propagate along the metal means that the electric field can lag slightly behind the electromagnetic wave, an effect that grows with the length of the wire, but the lag  is immeasurably small at the scale of a vacuum cleaner cord. The vacuum is switched on and the wave of energy, which we can think of as an electromagnetic field signal, travels at near light speed down the cord. All the mobile electrons in the copper wires immediately get the signal to start oscillating along the electrical circuit. The electric signal, as an electromagnetic (EM) wave, is composed of oscillating electric and magnetic fields. The electromagnetic energy flows as oscillating electric and magnetic fields that are generating just ahead of it. The electrons, acting as moving charge carriers and magnetic diploes, generate the oscillations of the (EM) signal. 

 

The EM wave loses energy along the cord, as energy is transferred into work done by the charge carriers in the wire. Energy is always lost during the transmission of electricity. Using Ohm’s law, we know that voltage, current and resistance are related to each other. Losses square with the current, which means that a small jump in current through a wire leads to a big jump in loss (through increased electrical resistance). This is why, for example, long-distance transmission lines are so high voltage, to minimize current loss due to resistance. A lot of energy is lost as microscopic friction (electrons bumping into electrons) inside the wire, and it is released as heat. A current-carrying copper wire warms up. Because the atoms are bound tightly together and there are many electrons in close proximity to one another, electrons will experience some resistance even in a good conductor such as copper.

 

Bloch Theorem

 

We haven’t explored yet how the three-dimensional lattice arrangement of atoms in a metal impacts its conductivity. In 1928, Felix Bloch formulated his Bloch theorem to deal with these effects. 

To use this tool to understand how current propagates within a metal lattice, we need to deepen our understanding of quantum mechanics once again. Electrons, like any matter particles, have a dual wave and particle nature. Resistivity can be explained by electron-electron collisions, a particle-like phenomenon. Conductivity, however, is better understood as the transmission of energy waves between electrons. Bloch theorem incorporates both natures simultaneously by treating electrons mathematically as wavefunctions. Put more precisely, this theorem is set up in quantum mechanics as special limitations put on Schrodinger’s famous equation. This equation describes the wave function of a quantum mechanical system. A wave function, in turn, is a mathematical description of an isolated quantum state. The wavefunction is useful because it gives us measurable information about that quantum state. I think it’s fair to say that the wavefunction, a complex function of space and time, pins down the un-pin-able physical properties of the electron. These physical properties are position, momentum, energy and angular momentum, the four quantum numbers mentioned previously. The “pin” is mathematical. If you know these four numbers, you know everything there is to know about any particular individual electron. In a quantum mechanical system the set of possible values for these physical properties cannot be specific values as they would be in a classical system. Here, they are expressed instead as eigenvalues, or as ranges of possible values.

 

You will notice I said “isolated state” when we really want to describe what is happening in a complex system containing many electrons. None of these electrons are acting as perfectly isolated particles, but again we start by simplifying matters.  We make the assumption that the electrons are acting like free particles and we ignore their interactions with each other. Copper, as mentioned earlier, comes pretty close to this hypothetical ideal state. An electron acting like a free electron in a metal such as copper can be treated mathematically as a wavefunction. In this case, we put limits on the wavefunction, so that our solutions to Schrodinger’s wave equation take into account the periodic nature of a three-dimensional lattice. Put more precisely, these solutions form a plane wave that is modulated by a periodic function. A plane wave is set up as a set of two-dimensional wavefronts traveling in forward through three-dimensional space in a perpendicular direction. A periodic function is a mathematical way to describe a system that repeats its values at regular intervals, like the regular intervals of a crystalline arrangement in a metal. These functions are called Bloch functions and we can use them to describe the special wavefunctions, the special quantum states, of electrons in a metallic crystalline solid. The Bloch function itself is not periodic but its probability wavefunction includes the periodicity of the lattice, which tells us that the probability of finding an electron is the same at any equivalent position in the lattice.

 

This mathematical description of electrons underlies the band theory of electron conduction that we discussed. Band gaps, energy states where electrons are forbidden, can now be described mathematically as values for energy, E, where there are no eigenfunctions in the Schrodinger equation. An eigenfunction, also known as a Bloch function here, is the mathematical description of any particular electron, as a wavefunction within a crystalline (metal) solid. We can predict the energy range where electrons are forbidden (the band gap) by working out where there are no eigenfunctions (no wavefunction solutions). The actual calculations are exceedingly complex and I don’t pretend to understand them. I do think, though, that it’s quite astounding that we have a way to precisely predict and describe a very complex behaviour that is critical to understanding how electric conduction works.

 

Emergent Behaviours and Properties

 

It might strike us as surprising that the Bloch theorem actually works so well, considering that we are ignoring many complications that could arise from electron-electron interactions. Complex interactions between subatomic particles in a metal can lead to the emergence of unexpected properties at the larger macro- or everyday scale.

 

The electron, as a elementary particle, has a charge and a mass. But, because it is a quantum particle, its charge and mass are qualities that can act in independent and unexpected ways. This is where our notion of the electron as a tiny physical ball of charge is really challenged. The Bloch theorem works because the electron’s charge moves within a periodic electrical potential as if it were a free electron in a vacuum. It’s mass, in contrast, becomes an effective mass. The mass of an electron at rest is always about 9 x 10-31 kg. Inside a metal, however, an electron can seem to have a different mass based on how it reacts to various forces acted upon it. Effective mass can range considerably, from zero to around 1000 times the electron rest mass, and this can have a big influence on how the metal behaves. For example, “heavy fermion“ metals, those with an effective electron mass in the 1000 range, can exhibit superconducting properties, such as zero electrical resistance below a low critical temperature. Effective mass can even be negative as in the case of P-semiconductor electron holes mentioned earlier. An electron hole is a lack of an electron mass where one should exist in the lattice, and leaving a local net positive charge. Each hole acts like a particle and is referred to a quasiparticle, in this case a positively charged one. It is phenomenon that arises from a complex system. This behaviour plays an important role in current conduction through semiconductors. Excited electrons leave behind holes in their old ground state energy level, and they can move just like electrons do, resulting in an electric current moving in the opposite direction.

 

Conclusion

 

Electric current is perhaps the most basic concept at the heart of the science of electricity. It’s a rare day that goes by when we don’t make use of at least one light bulb or our mobile phone. It’s so familiar to us that we might not give it much thought. Yet, electrical current is mysterious. It cannot be directly seen, heard or felt. It’s not easy to gain an idea of what it actually is. In order to do this we had to dive deeper and deeper into theory, from classical to quantum mechanical, while we updated our mental snapshot of the electron along the way, from the Rutherford haze to the tiny charged billiard ball to a modern quantum cloud of mathematical probabilities.

 

 

 

Saturday, February 1, 2020

Time as a Dimension

To explore this, we are really exploring the model that incorporates time as a dimension. We have no doubt heard of space-time, which we might imagine as some kind of four-dimensional fabric that permeates the universe, or as something that sets the stage upon which the universe exists. We might think of Albert Einstein as the father of space-time, and though several physicists took starring roles in developing this theory, Einstein no doubt brought the theory together.

All events in the universe take place in space-time. Space-time is actually a mathematical model. It fuses three spatial dimensions with one time dimension into a geometric whole. How can we conceptualize this? There is a common trap in physics, which is to confuse the theory or mathematical model with the actual thing, something physical that can be observed and measured. A wealth of experimental evidence, listed in the 4-minute video below, supports the theory of special relativity, which describes the space-time model:



This is good evidence that the space-time model used in special relativity accurately accounts for real observable and testable phenomena.

What is a Dimension?

Mathematically speaking, a dimension is the smallest number of coordinates you need to specify the location of a point. A line, for example, has a dimension of 1 because you only need one coordinate to specify a point on it. A surface or plane has a dimension of two and the inside of a sphere has three dimensions. To describe the location of a point inside a sphere, for example, you need three coordinates - we usually call them x, y and z coordinates. Building on this we can say that a point within a sphere (or any three-dimensional space) can potentially move in any combination of three possible directions - in the x, y and/or z direction. Borrowing from statistics, we can say that point possesses three degrees of freedom.

It's easy to visualize the three dimensions of space we live in. We've got up/down, latitude and longitude. An everyday example of a location in three-dimensional space might be "on the third floor in the northwest corner of the Black Building." We can pinpoint exactly where to go if we are given these directions. In our everyday world, time, however, feels different. Locations in space can be fixed but we experience time as fluid. It flows as events constantly recede into our past and reach into our future. Now when we revisit our directions, we notice that we need a "when." What if these are instructions for a dentist appointment, for example? The dentist wants me at a specific location at a specific time, such as in his chair at 2 pm on Wednesday. I've got all four coordinates you need. At 1 pm or 3 pm, someone else will be occupying that chair, those identical spatial coordinates. But only I will be in that chair at this specific time coordinate of 2 pm Wednesday. Therefore, we can see that to describe any event we'll need both spatial coordinates and a time coordinate.

But how is time related to space? I can explore this by describing my walking path to the dentist's chair. I will need to describe both space and time. I enter the east door of Black Building at 1:45, approach the elevators at 1:50, wait there for one for a minute for the door to open and then travel up to the third floor during the next minute before I walk down a corridor to the dentist office at 1:55. I plunk myself down in the dentist chair at 2 pm and remain there for 50 minutes. I am describing a series of events that flow through space AND time.

During these events spanning between 1:45 and 2:50 pm, I am moving through space and time, in other words, I am moving through four changing coordinates, or dimensions. However, I notice that I may be moving through one or two spatial dimensions simultaneously but in every case as I move through space I must also move through time. Time seems to have some rules that space doesn't. I can never "stop" in a time dimension. I can never move into a negative time dimension. Time, then, appears to have fewer degrees of freedom than any spatial coordinate. I, or anyone, can only move in one direction along this timeline and at a rate I can't control. Is time a dimension? If it is, it does not appear to be a fourth dimension in the same way as the other three spatial dimensions.

We can map out my progress between 1:45 and 2:50 pm on Wednesday using a framework that describes changes in all four coordinates of time (1) and space (3). Yet, as we just noticed we encounter some problems when we think about time as a fourth dimension equivalent to a spatial dimension. We can describe this impasse mathematically. As we will get into later, it was Hermann Minkowski, not Einstein, who formulated the dimensions of space-time, and these four dimensions are NOT equivalent to four-dimensional (Euclidean) space. Space-time is not, for example, this albeit mesmerizing four-dimensional rotating cube:

JasonHise;Wikipedia
This cube above is described by Euclidean space. There is no time dimension to it. Space-time does have geometry but it is different in important ways. The non-Euclidean mathematics used to describe space-time describes how time works with space. To make this connection conceptually, we can trace how the concept of space-time came about.

Time Was Once an External Stage

Isaac Newton: We probably all started our exploration of physics with this great man, a key figure in the scientific revolution in mid-seventeenth century Europe. In his time, a four-dimensional framework for any physical event would seem unnecessary and probably ridiculous. To describe his laws of motion, he needed only three spatial dimensions, the ones we experience every day, and he assumed it all happened while an absolute time progressed at a specific rate independent of everything else going on. Even physical space was likewise treated as outside all events. Every object either had an absolute state of motion or an absolute state of rest relative to the absolute space it found itself in. Time and space were treated much like an external stage on which all phenomena in the universe take place. The assumption made sense and it worked for a long time. It is how we experience space and time every day.

The idea was basically unquestioned until the mid-nineteenth century, shortly after James Clerk Maxwell and others, started to tinker with electricity, magnetism and light. Maxwell discovered that all three were a) related to each other and b) traveled as disturbances through (Newtonian three-dimensional) space.

Maxwell worked out that "light and magnetism are affectations of the same substance and that light is an electromagnetic disturbance propagated through the field according to electromagnetic laws." The "field" in this statement is the epicentre of what would become one of the most hotly contested debates in the history of science. What exactly was this field? His work would ultimately lead to the connection between space and time. To start with, he and others knew only two well-established facts: 1) light seemed to have a constant velocity through air and 2) light slowed down when it traveled through other transparent media such as water.

Aether Was the Medium of Space

Scientists at this time knew that electromagnetic disturbances act like waves so they figured they must travel through some kind of medium, which they called aether. Experimental results on aether, however, were confounding. For example, the speed of light through air was constant in any direction. Wind or air density had no effect. Even more confusing, the speed of an electromagnetic disturbance seemed to be independent from speed of the source of that disturbance. In other words, light seemed to disregard Newtonian physics. How? What about this aether substance could accommodate such findings? It proved frustratingly impossible to define electromagnetic waves mechanically. They did not act like other mechanical waves such as sound waves or water waves.

Like many other scientists at the time, Einstein, pondered how aether worked. Like physicists Paul Dirac, Louis de Broglie, Maxwell and others, he thought that aether was some kind of medium with physical properties filling empty space. There must be something there that carries the electromagnetic disturbance, and perhaps the results could be explained by some kind of elastic force through which the waves are propelled, analogous to water waves or sound waves. Even in 1920, after he developed special relativity, he stated that there must be something that allows for the "existence of standards of space and time (measuring-rods and clocks)" to allow for space-time intervals in the physical sense.

Aether Theories Run Into Trouble

Various aether experiments designed to find out how it worked mechanically, and there were many, yielded either contradictory or nonsensical results. An interesting and well-known example of this problem was the Fizeau experiment, conducted in 1851. At that time, a number of scientists were comparing the speed of light (an electromagnetic disturbance) in air versus water. They could observe that a beam of light slowed down in water. They wondered what process slowed the light-bearing aether down in the denser medium.

If you are wondering how light slows down when it has one invariant velocity, you ask a fantastic, often overlooked, question. A good concise answer can found here.

This phenomenon, called refraction, has been observed since ancient times and was described mathematically in the 1600's by Snell's Law. A number of researchers revisited refraction as possible evidence that aether could be partially dragged by moving matter such as water. The idea being tested was that aether moving against the direction of water flow might be slowed down. Vice versa, aether dragged along with the water flow might boost the flow rate of the aether and therefore the speed of light. Fizeau designed an experiment that compared the speed of light through water moving in the same direction as the light beam with the light's speed moving against the direction of the water flow. If their theory was correct, light would move faster along the same direction of water flow and slower when it's against the flow. They didn't know how or if aether interacted with matter but this experiment was designed as a first step to the answers. They made the assumption that because light could penetrate all transparent media, such as water, those media were permeable to the aether. If the medium is moving, does it carry the aether along with it? Is the aether partially dragged or is not affected at all?

Hippolyte Fizaeu got puzzling, but not entirely unexpected, results from his experiment. He showed that the speed of light in same-flow water was boosted but it was less than the sum of the speed of light in air plus the speed of the water flow, as Newton's laws would have predicted. The aether appeared to be dragged along, but only partially. It turned out that Augustin-Jean Fresnel had already established a dragging coefficient in the late 1700's, based on several earlier aether experiments that appeared to support the idea of partial aether dragging. All of these experimental results seemed to suggest that the aether might be denser inside mediums such as water than it was in air or in a vacuum and that light traveled more slowly through denser media. Fresnel's dragging coefficient was proportional to the refractive index of the medium.

Scientists at the time knew that the reduction in the speed of a beam of light depends on the index of refraction, which in turn depends on the light's wavelength. The refractive index decreases with increasing wavelength, so, for example, blue light bends more than red light when a light beam passes from air into water. This is why white light disperses into a rainbow when it passes through a prism. This presented a snag. The experimental data pointed to a seemingly complex scenario where aether is partially dragged by matter and the aether must flow (simultaneously!) at different rates for different colours of light, as a white light beam, containing all the colours, was used in their experiments. Were there a seemingly infinite number of aethers, one for each wavelength of light? The use of polarized light in this experiment presented the same problem. Light polarized in opposite directions both exhibited the same partial aether drag, suggesting that the aether was carrying two opposite directions of motion at the same time. These partial-dragging results were confirmed by numerous other experiments as well. Aether theory was offering complication rather than simplification.

Electromagnetism Re-imagined

It wasn't until around 1892 that Hendrik Lorentz approached these baffling experimental results from a new angle and a solution began to take shape. He assumed, first of all, that the aether was completely stationary.  He was then able to derive Fresnel's coefficient by using Maxwell's equations and an undragged aether. He looked at the problem as one of light speed transitioning between two reference frames, one where the system is at rest in the aether and the other where the system is in motion in the aether. By rest, he meant absolute rest in the absolute space of Newton. By doing this he introduced a clear distinction between matter (this time in the form of electrons) and aether. This meant a departure from any mechanical theory of aether.

Referring the work of Maxwell and others, he described the aether as "states" in an electromagnetic field. By doing so, he introduced an abstract aether replacing the previous and problematic mechanistic model. He also planted the seed for special relativity: A moving observer with respect to the aether will observe the same electromagnetic phenomena as an observer at rest in the stationary aether.

Lorentz's interpretation was that partial dragging was something that happened to the electromagnetic wave itself and not something that happened to the aether, which was stationary. This proved to be a critical first step in the evolution toward our modern theory of space-time. It was, however, a first step. It transformed mechanical aether into an abstract electromagnetic aether, but it still held onto the presence of some kind of aether and it held onto Newton's concept of absolute space.

The idea that space must contain something to support the propagation of electromagnetism (and gravity as well) was extremely difficult to let go. By 1901, as Lorentz was developing the theory underlying his famous Lorentz transformation, which would provide a bedrock for Einstein's theory of special relativity, Henri Poincaré wrote (in a state of philosophical angst?) that there must be no absolute space nor absolute time. Even so, Poincaré would not let go of aether: "If light takes several years to reach us from a a distant star, it is no longer on the star, nor is it on the Earth. It must be somewhere, and supported, so to speak, by some material agency."

Aether Evolves into Space-time and Fields

Quoting from Einstein once again in 1920, after he published his theory of relativity, "To deny the aether is ultimately to assume that empty space has no physical qualities whatever. The fundamental facts of mechanics do not harmonize with this view." Again, one can detect the angst.

Today, we can argue that Poincaré and Einstein were correct about space, but only in a sense. We can argue that the field itself has replaced the aether. Electromagnetic radiation, such as visible light, which was the focus of Fizeau's and other aether experiments, is carried as waves through an electromagnetic field. These waves operate by quantum rules rather than mechanical rules. We now have a concept of a physical universe that is described by the mathematics of space-time geometry, and within that geometry we describe various fields carried by force-carrying particles called bosons. Photons are the force-carrying particles of the electromagnetic field. Starting with the development of the theory of electromagnetism in the late 1800's and continuing through the development of quantum mechanics later in the early 1900's, physicists gradually abandoned the idea of a background medium altogether. Aether isn't required to explain how special relativity works.

How "real" is a field? Is it physical or strictly a mathematical construct? The concept of the field arose as a fundamental physical quantity that independently exists. For example, physicists envisioned the electromagnetic field extending indefinitely throughout space in all directions as a physical field interacting with matter. Maxwell's electromagnetic field equations were developed using classical field theory which obeyed Newton's laws. Later on they were refined further by incorporating special relativity and quantum mechanics. We now know that an electromagnetic field is carried by force-carrying photons. Subatomic particles, such as photons, are treated as excited states in the quantum field that obey the laws of quantum mechanics.

As treated in quantum field theory, a field is strictly mathematical and doesn't physically exist. The field still extends all over space and we can make an observation of the field by taking a measurement of it at a particular moment and location. In this sense, the field interacts with matter and we can argue that the field is physically real. We experience evidence of it all the time, such as when we see a burst of visible light photons when we turn on a light. We feel a static charge or watch iron filings arrange themselves according to the lines of force exerted by a magnet. In the case of magnetism, for example, we indirectly detect the virtual photons that carry the magnetic force. We understand these photons as mathematical wave functions.

The Speed Of Light Forces a Shift In Thinking

Einstein, wondering about space and time and perplexed by the nature of aether, turned his focus to the experimental findings. He knew that Lorentz was beginning to approach the aether problem in terms of changing frames of reference. He focused on one blaring observation. The speed of light appeared is a universally unchanging value. In itself this was a truly mind-blowing observation in a then largely Newtonian universe.

Imagine the headlight of a train approaching at the speed of light. The photons in that light beam would also be travelling at the speed of light, and never exceeding it. Why don't these two velocities add together, as they would if a man threw a ball forward from a train travelling forward at everyday speed? If photons followed the same rules of Newtonian dynamics as balls do, the two speeds would add up. What is it that slows the light down and keeps it in check? What is that process if it is not the work of some kind of elastic aether? Einstein knew that something in the description of this thought-experiment must give. One thing he could conclude with some certainty was that he did not yet have the entire picture of space as the medium through which light travels. By following a tactic similar to Lorentz by allowing the question of medium to take a back seat, he could reframe the problem. If the speed of light never changes, then time or space, or both, must. Put mathematically, space and/or time must transform.

Time Can Vary

The concept of transformation itself isn't new in physics. Galilean transformations operate in Newtonian physics. They tell us that any event that takes place in one frame of reference will operate under the same physical laws if it takes place in a different frame of reference. For example, barring all other sight cues, a car traveling at 50 km/h passing a car traveling at 30 km/h in the same direction will appear to the passengers of the 30 km/h car to be traveling at 20 km/h. It's the basic addition/subtraction of velocity vectors, operating under the same rules as the ball being thrown from a train example above. These kinds of transformations presume that the passage of time is the same for observers in different frames of reference. They presume that time is absolute in other words. These Newtonian rules are still useful and that's why we learn them. They work perfectly until we are dealing with velocities approaching light speed (or near gravitational fields). Lorentz and others tried to understand how the speed of light breaks these well-established common-sense rules. In any reference frame the speed of light is always the same. It does not obey the Newtonian laws that underlie a Galilean transformation.

A Galilean transformation holds up for events that happen at everyday velocities, but as an object approaches the speed of light in one reference frame compared to a stationary reference frame (we can call this frame a stationary observer), both space and time, for that observer, transform. Space and time depend on the reference frame. Any object approaching the speed of light experiences time dilation (time stretching or slowing down) and length contraction as observed relative to a stationary observer. To that observer, the object contracts in the direction it is traveling* and a clock attached to that object slows down.

*An object travelling near light speed will actually appear rotated even though its measured length will be contracted. The object is moving so fast that light from the along the object reaches the observer at slightly different times. A receding object will appear contracted and an approaching object will appear elongated, while a passing object will appear skewed or twisted. This optical effect is called Terrell rotation).

This means that observers moving at different speeds relative to one another can observe different distances, different elapsed times and even, as a result of these transformations, experience different orderings of events. These transformations are not illusions. At the expense of getting ahead of myself, consider an example of a proton (a particle of matter) in the Large hadron Collider. It is accelerated to almost light speed and as it does so it experiences a Lorentz factor of about 10,000. The Lorentz factor is the factor by which time and length change for an object that is moving. To put this in perspective, if you could shrink down and ride on top of this proton from Earth to Alpha Centauri, your trip would you take only a couple of days. Alpha Centauri is four light-years away, which means it takes (traveling at light speed!) four years for its photons to make that same length of trip. An observer on Earth would record that your trip to Alpha Centauri took a little over four years.

Putting his thought-experiment observations into a formal framework, Albert Einstein published his game-changing theory of special relativity in 1905. It incorporated Lorentz transformations in space and time. A few years later, Hermann Minkowski formulated a geometric interpretation of the Lorentz transformations, and this is now the mathematical structure, called Minkowski space-time, on which the theory of special relativity rests.

Minkowski space-time mathematically combines three-dimensional Euclidean space with time to create a four-dimensional structure called a manifold. A manifold is a strictly mathematical concept that is nicely explained here.

The Space-time Interval

If we take the simple concept of distance, we can get a feel for how Minkoswki space-time works. In Newtonian physics, the distance between two points is invariant. It will be the same regardless of reference frame. In special relativity, however, that distance will depend on whether the observer is moving or not (length contraction). In four-dimensional space-time, a new invariant "yardstick" called the space-time interval replaces distance. Whereas in Newtonian physics, time and space (distance) are invariant, in Minkowski space-time, time dilates as distance contracts. These measurements are dependent on the frame of reference. The space-time interval of an event, which combines space and time, is the same in any frame of reference. A space-time interval extends from one place and time to another place and time. We can even build space-time by taking successive snapshots of space over time and adding them all together. Measurements of space and time can vary between observers but the space-time interval, obtained by measuring the distance and time between two events, will always be the same in every frame of reference. It doesn't matter how fast or in what direction an object is traveling with respect to the observer. The space-time interval displays Lorentz invariance.

To visualize how time and distance relate to one another geometrically in a space-time interval, as well as how the speed of light is constant in every frame of reference, try this 7-minute video:



How To Get From Length to a Space-time Interval

How did we get to this new invariant idea of "length" in space-time? When Minkowski developed the space-time manifold, he imagined that the space-time interval could be related to Pythagoras' theorem, but in four dimensions rather than the two we are all probably familiar with when we draw a right triangle on a sheet of paper, shown below right.

The Pythagorean theorem states that the length of the hypotenuse, z, is given by the square root of x2 + y2 where x is the horizontal measurement and y is the vertical measurement in a two-dimensional coordinate system such as a sheet of graph paper. If we want to describe z's length in three dimensions, we just add a measurement along an additional horizontal axis, w, which we can imagine as a line coming out of the page. Then we get z2 = w2 + x2 + y2. We can now describe line z's length and position in 3-dimensional space. To measure z's coordinates in time as well as in space, Minkowski introduced a time dimension, (ct) to the equation. Here, c is a conversion constant, which is the speed of light in a vacuum (metres per second), and t is the time interval (seconds) spanned by the space-time interval. We can think of it as the distance light travels in t seconds. This is a way to incorporate a new "length" along a new axis, and as we do this we are switching to a four-dimensional coordinate system. By doing so we are bringing time, as a unique dimension, into our geometry. The speed of light conversion constant makes this dimension uniquely different from the other three spatial dimensions. It also tells us that the speed of light can be used as an invariant measurement of time called proper time.

With a little mathematical finesse, we end up with three dimensions of space and one dimension of time:  s2 = w2 + x2 + y2 + (ict)2. We've changed our variable z to s, to show that we are now measuring distance as a space-time interval. We've also added a new variable, i, to our time axis. The term i is an imaginary unit, also known as √ (-1). The old-fashioned picturesque descriptor "imaginary" doesn't mean an imaginary number is made up. It just helps us find solutions to mathematical problems. In our case, imaginary time is real time that undergoes a mathematical transformation called a Wick rotation. A Wick rotation is a way to convert a problem in Euclidean four-dimensional space into a problem in Minkowskian four-dimensional space-time. It trades one spatial dimension for a time dimension and allows the dimension to undergo a Lorentz transformation, which mathematically is a rotation of coordinates.

Since the ict term is squared we can multiply (ct) by -1. We end up with s2 = w2 + x2 + y2 - (ct)2. Again, we can think of this (ct)2 variable as "distance" along the time axis.

The introduction of an imaginary unit hints to us that even though we've put our equation into the form of a Pythagorean equation, the time dimension in it doesn't "act" like the other spatial dimensions. It does not have a simple Euclidean geometrical relationship with space.

Physicists now describe space-time in terms of a newer mathematical construct called a metric tensor. General relativity also describes space-time, but in that case, the space-time needs to curve under gravity. We can think of the metric tensor is a device that makes corrections to Pythagoras' theorem to enable the right triangle we used as our starting point to map onto curved space-time. It also does away with the imaginary unit (i) we discussed earlier by describing events in real time instead. The negative sign, however, is preserved in the metric tensor but it now describes how distance changes with time when space-time curves. The original Minkowski equation describes an incorrect but simplified flat space-time.

By creating a space-time interval, we can understand both the invariance of the speed of light as well as time dilation and length contraction.

Speed of Light Invariance

Imagine a photon traveling at the speed of light, c. All observers will observe that same velocity no matter what their velocity might be relative to it. The distance traveled by the photon (let's say it's traveling in the x direction so we'll call it distance x) for t seconds can be written as:

x = vt  where x is the distance traveled, v is velocity and t is time

Let's begin to transform this simple equation for distance into one for a space-time interval. First we'll incorporate it into the Pythagorean theorem in three dimensions, like we did earlier. I'll start using s for the distance even though I'm not quite correct yet because we haven't incorporated time.

s2 = x2 + y2 + w2

There is no motion in any direction except the x direction so the w and y axes are zero. We need to describe this relationship in terms of space-time so we add the time dimension [-(ct)2]. We can now properly describe the distance (x) in terms of a space-time interval (s):

s2 = x2 + 02 + 02 - (ct)2

We can swap out x by incorporating our earlier equation x = vt.

s2 = (vt)2 - (ct)2. Our object is traveling at the speed of light so we know v = c.

 s2 = (ct)2 - (ct)2. We get s2 = 0 so s = 0.

This means that the space-time interval for any object traveling at light speed is zero. It is invariant. It doesn't matter what reference point you measure the object from. You could be accelerating in the w or y direction as you measure its velocity. It will always be light speed and its space-time interval will always be zero. Put another way, only an object traveling at light speed will have a zero space-time interval. All observers will observe the same (zero) space-time interval for that object, which means they all observe it to have a velocity of c.

How does a photon experience the universe? We can get a feel for this surprisingly complex situation by comparing the world lines of three objects, all traveling at different constant velocities in the same direction, shown below in a simple space-time graph.

Jheise;Wikipedia
We have to be careful because we are not representing velocity or position versus time. Each line, called a world line, is built from a sequence of space-time events for each object. Each point on each line is a four-dimensional space-time event. An event in space-time is a specific location in three-dimensional space at a specific time. The t in the graph depicts proper time.

In this graph, t is time and x is distance along one space coordinate. We could draw a more complex space-time graph by incorporating all three space coordinates with one time coordinate, representing Minkowski space-time.

An object at rest would be a vertical line originating at the same origin point as the coloured lines and where x = 0. Its world line is space-like. A space-like world line could likewise describe the length of a physical object such as a ruler, as the distance between two space-like events. Each of the three coloured lines represents the world line of an object traveling at a specific constant velocity (hence all the straight lines). Their world lines, and the world lines of any objects traveling less than the speed of light, are time-like curves in space-time. Even though only straight lines are drawn here, any world line is considered to be a special type of curve in space-time.

This graph represents all times (future and past) and all possible distances along x in space-time. It is a simple representation because all three objects are traveling along in the same direction, along the x-axis. They all originate at the origin of time and distance on the graph. At that origin point, they share the same space-time interval. A physical example might be a single particle decaying into three particles, each having a different velocity.

A stationary object moves in time but not in distance. A slow object moves further in time than it does in distance. A faster object moves further in distance than it does in time. A very fast object moves in distance but very little in time. A photon has the fastest possible velocity. It moves in distance but not in time. It follows a light-like curve, which would be represented as a horizontal line moving along the x-coordinate, where t = 0. A light-like curve is a straight line in this simple graph where two spatial dimensions are not shown. Often the convention is to draw this horizontal line at a fixed 45-degree angle. By doing this we can draw the light-like curve in three spatial dimensions as an easier-to-visualize three-dimensional cone, directed upward into the future and downward into the past, shown below.

MissMJ;Wikipedia
Put more sophisticatedly, we can say the photon approaches the limit of proper time. A photon might be emitted from a distant star and travel through space for 4 billion years before it is absorbed. We can measure that journey as taking 4 billion years (a great "distance" in time), but for the photon there is no distance along the time axis. It is emitted and absorbed instantaneously. Remember Henri Poincaré's troubling question? He asked where the light beam is in the space between stars. He wondered what medium was carrying it. Can we say that the photon even has a journey? We observe photons of starlight traveling across great distances over light-years of time. But in the photon's frame of reference the universe does not seem to consist of the space-time we experience.

Time Dilation

Imagine setting up an array of synchronized clocks over a very large table in space, a table on the scale of thousands of kilometres across with no gravitational field are nearby. From one edge of the vast table you take a photo of all the clocks. You find that the clock closest to you is running a little faster than those furthest away. After a little thinking, you realize you need to take into account the transition time for the light from each clock to reach you. The speed of light is constant so this is a fairly straightforward synchronization calculation. You go back and adjust all of your clocks. Now they all read the same time when you take your photo of them. A friend flies past your clock arrangement at 0.95% light speed and takes his own photo of the array at exactly the same time you take your next photo. Comparing photos, you notice that his clocks are all a bit behind yours. Your frame of reference is at rest compared to your clock table so you don't experience time dilation. However, you and your clock table are in motion compared to his frame of reference. He records time dilation. His present moment was not the same as your present moment. The two events you experienced were desynchronized.

If an additional initially synchronized clock were glued to the outside hull of your friend's ship beforehand and you repeated the experiment, you might guess that you will see it as running faster than your clocks. Instead, you read it as running slower as he flies past you. For you, he is the moving frame of reference. For him, once again you are the moving frame of reference. You once again see in your photos that his clocks are slower. And yet, for him, your clocks were slower. This counter-intuitive effect is known as the twin paradox.

What is different in the two frames of reference has nothing to do with the mechanisms of the clocks. A moving mechanism doesn't get heavier or something such as that. The key difference is that the moving clock is traversing a longer distance between events (ticks). The events do not have to be hands moving on a clock face. The clocks could be mechanical, quartz digital, atomic or even hourglasses.

To get a feel for this it might be easier to imagine that our clocks are made of pulses of light bouncing between two mirrors. One trip from mirror to mirror is equivalent to one tick of our earlier clock. The speed of light is invariant, so the ticking mechanism of this clock will be perfectly constant. The clock moving with respect to the stationary clock will tick slower. The moving clock will, in its frame of reference, experience the stationary clock as the one moving and it will tick slower than the former one. This brings home the fact embedded in special relativity that there is no absolute motion and there is no absolute rest. The only absolute is the speed of light. We can visualize time dilation (and the twin paradox) in the set-up in the gif below.

Cleoris;Wikipedia
Each blue dot represents a pulse of light. Each pair of dots (red pairs and green pairs) are mirrors bouncing light pulses back and forth. Each pair is a clock. If we measure the time it takes for one light pulse to reach from one mirror to the other mirror we will always get the same result as long as we are in the same frame of reference. This is called proper time. It is the fastest possible time because it is the shortest distance traveled by the light pulses. For each group of clocks the other group ticks more slowly because the light pulse has a longer distance to travel when it is moving. Time dilation works not just for light. It works for any series of events. The physics of any event or process is constrained by relativity. In other words, special relativity forces all other laws in physics to obey. For example, a man in motion with respect to a stationary man ages more slowly.

Length Contraction

We can do another thought experiment to explore how length contraction occurs. Imagine two light clocks like the ones described above in which light bounces between two mirrors. We can put them close together or far apart and we can orient them in any way we want with respect to each other. If they are both at rest with respect to us, as observers, they will run at the same rate. Any direction or location in space (barring any and all gravitational influences) is physically the same and physical laws work the same anywhere in the universe. If we set them perpendicular to one another and then set them both in motion at 99 % light speed in the direction of one of the clocks, we get some interesting results.

We, as observers, remain at rest with respect to the clocks. This means that as the clocks fly by us, the light is bouncing parallel to the motion in one clock and the light is bouncing perpendicular to the motion in the other clock. We will find that they will both run slower, as we expect, but we also find that they are still both running at the same rate. This observation is not what we expected.

The clock that is perpendicular to the motion should slow down, as we figured out above. We can imagine that those light pulses must travel a longer distance because they are making a stretched out zigzag path of motion between the mirrors. The stretching is where the extra distance comes from. But what happens to the clock that is parallel to the motion? Aren't those light pulses going to take a much longer time to reach the front-facing mirror that's going just 1 % slower than them?

We can put together a more concrete example to show what's going on here. Let's say the mirrors in the clocks are 300,000 km apart, the distance light travels in one second. At rest with respect to us, the light will take one second to travel from one mirror to the other. Now we set them in motion like we did above, with one set of mirrors oriented perpendicular to the motion and one parallel to it. If the clocks are now moving at 99 % light speed and the light is moving perpendicular to the direction of motion in one of them, we can calculate that the light will now take about 10 times longer, or 10 s, to make one trip between mirrors. The clock is now going 10 times slower relative to us.

What about the parallel clock? At rest, the light takes one second to go from one mirror to the other. If the light pulses in the clock are moving in the same direction, at 99 % light speed, the light has to chase a rapidly receding mirror in one direction. We can figure out that it will take about 100 s to reach the front mirror when it's moving away at 99 % light speed. It will take just a tiny fraction of a second to make its return trip to the other mirror because the back mirror is approaching at almost light speed. We discover that the light takes ten times longer to bounce mirror to mirror in the parallel clock (100s rather than 10 s). Yet we measured them and they are both running at the same rate - 10 times slower than they did at rest with us. How is the parallel clock still keeping the same time as the perpendicular one? The only way it can is to physically shrink in the direction of the motion, shortening the bounce distance. In fact, it will shrink to 1/10th of its rest length at 99 % light speed. Length contraction and time dilation have a perfectly inverse relationship. Length and time compensate one another to preserve the invariance of the space-time interval we explored earlier.

If one of the clocks could travel at light speed (and it cannot because it has mass) the light pulses in it would not move at all. There would be no ticking forward in time. If we could somehow ride along with a photon of light, we would discover that it does not experience time. If we consider the effect of length contraction as well, we come to a startling conclusion. There is no distance at all between the two mirrors in the parallel clock. In that clock the mirrors themselves would have no depth. Time slows down and length contracts for objects travelling very fast. A photon's path of travel is shortened to zero. Proper distance, like proper time, does not exist for a photon. Realizing this gives new weight to the concept that light has a space-time interval of zero.

Some Parting Thoughts

To accept the well-established fact that we live inside time as part of four-dimensional space-time is a bit like accepting that we experience only a tiny sliver of visible colour within a far vaster array of electromagnetic radiation. We know that it exists but we don't directly experience it in our everyday lives. We intuitively understand the universe in three-dimensional space but it is almost impossible to conceptualize four-dimensional space-time.

The movie Interstellar plays with the fact that time is a dimension. In that movie, a future "us" has figured out how to manipulate time to make it act like a tangible physical dimension. We currently can't do that and I don't know if that ever could be possible, but the mathematical formulation of space-time, in particular the Wick and Lorentz rotations, appear to treat time and space as two facets of the same thing. We got to our understanding of time as a dimension of space-time by way of the speed of light. At the speed of light, both time and space reach their limits. Does space-time exist at the speed of light? Do space and time fully unfold to our perception only when we experience the universe at rest?