This week’s essay was going to be about “neurodivergence”, the phenomenon in which differences in neural “wiring” between people can set the stage for marked behavioral differences from other people. The idea was to explain the term, discuss how wide-spread it is, and then explore its genetic basis, probably in several installments. However, thinking more about this, I decided that the topic required some background genetic information before entering this topic. That is what this piece will provide.
In particular, I want to explain what I think is a common misperception about the genetic basis of complex traits, physical and behavioral, in health and disease. That misunderstanding involves three elements. The first is the belief that if the foundation of some biological problem is genetic, it involves one or a small number of genes that “determine” it. This is unstated but a common popular assumption. The second is a corollary: identify those genes and you are half-way to curing or ameliorating those conditions. (If the small-gene-number assumption is false, however, the corollary must be also.) In fact, most traits are determined by many, many genes, a fact denoted as “polygenic inheritance:” The third element is more of an omission than a mistake: ignoring the immensely important matter of having the “right” quantitative expression of genes. If there are differences, even minor shortfalls, in certain gene “activities”, there will be biological changes in either how the individual develops or their physiology works.
These points may seem too technical to be interesting but appreciating them is vital for understanding the basis of complex traits and how they need to be approached. First, however, we need to define four key terms: A) “gene”, B) “gene product”, C) “gene activity”, and D) “mutation”.
A “gene” (A) is a stretch of deoxyribonucleic acid or DNA, the double-stranded large molecule that carries the hereditary information in our chromosomes, that contributes an element of heredity . The idea of the gene as the “unit of heredity” seems simple but is not; see no. 5, “On how the term ‘the gene’ first gained, then lost its meaning”. Nevertheless, it will serve as shorthand to indicate an element of heredity consisting of DNA.
The term “gene product” (B) only rarely appears in popular accounts but is key to understanding how genes work. Genes do not directly do things but act through molecules made from them; these are all in the class of ribonucleic acid or RNA. RNA is a close chemical relative of DNA, though single-stranded and a copy of one of the two complementary DNA strands of a gene.1 Most RNAs copied from genes move out of the nucleus, where the chromosomes are contained, into the surrounding semi-colloidal “cytoplasm” of the cell. There are said to be “messenger RNAs” and they are “translated” into single proteinaceous molecules called “polypeptide chains”; these gene products are the molecules that do most of the work of the cell.2
There is, however, another large class of genes copied into RNAs that do not specify proteins; these RNAs do their work directly. They act mostly in the nucleus, there to “regulate” other genes (increase or decrease their RNA production) though some also act in the cytoplasm.3 I will devote a future article to them but here will focus on the first class, the “protein-coding” genes.
The third term, “gene activity” (C), is also hardly ever clearly defined but is essential. It denotes both the specific qualities of the gene product that give it its characteristic biochemical function in the cells in which it appears AND its total quantity. The crucial point is that a particular gene product, with its intrinsic biochemical capability, is needed in certain amounts for normal functionality – hence at a standard “activity” in the cells/tissues in which the gene acts. Each such cell or tissue type has its own requisite level for that gene activity, hence a gene does not have one level of activity but multiple ones for the multiple cell types in which it acts; that versatility characterizes most genes.
Finally in this list of basic terms is “mutation”. “Mutation” is the blanket term for new hereditary changes in genes. Genetic changes in the DNA that decrease (or, rarely, increase) a gene’s activity can cause changes in the operations of those cells and tissues. Those alterations may cause problems, including health problems, or at least differences from other people. (Such differences are at the heart of neurodivergence, as we will see.) If a new mutation takes place in the “germ line” – the cells that give rise to egg cells or “ova” in females, and sperm in males – and is passed to the next generation, it can become hereditary, namely transmitted to many successive generations. Such mutations are then referred to as “genetic variants”.
Mutations or genetic variants come in three broad classes but those groups are not of equal importance. The largest group are those that have little or zero effect. They are termed “neutral” mutations. How can a genetic change do nothing or very little? Simply by being within a stretch of DNA sequence that does nothing or very little. Such sequences comprise a large and unknown fraction of the human genome but (probably) well over 50%. This is a fascinating phenomenon: why do most complex organisms carry around so much seemingly useless DNA? The likely answer is that such sequences actually do something but not dependent on the details of the sequence (unlike the protein-coding sequences). I will write on this later but the key point here is that most newly arising genetic changes are neutral in their effect.
The second class of new genetic variants are those that either increase the amount of gene activity of one or more genes, or change its properties slightly. Such “hypermorphs” or “neomorphs” (apologies for more genetic jargon) are probably the smallest class of new mutations. Each is potentially important but they are comparatively rare.
It is the third class that is most important in terms of aggregate effect: these are the genetic variants that decrease the amount of gene activity. Such mutations are said to be “loss-of-function” mutations or, as abbreviated here, l-o-f mutations. The overwhelming preponderance of mutations that have an effect on the organism – whether creating a problem or just a difference – are l-o-f mutations. A crucial point is that such mutations can cause a complete loss of activity (so called “null mutatiions”) or only a partial one (“leaky” mutations).
Our collective understanding of these genetic basics has been evolving for roughly 125 years, since the birth of modern genetics in 1900.4 . As someone trained in genetics and who has always been interested in the history of ideas, I love this story but I know that not everyone does nor is there enough space to cover it. I will, however, sneak in a small fraction of this history.
What I want to do in this article chiefly is convey the importance of quantitative matters in understanding genetics. Specifically, I will discuss A. the importance of the correct amounts of different gene activities and B. the large number of genes that participate in forming any complex trait or condition.
The foundational question of the field of genetics concerns the relationship between the genetic make-up of an individual – its “genotype” – underlying a trait and the way the trait appears – its “phenotype”.5 It was not resolved in the early days of the field and remains contentious though better understood today.
Nevertheless, a basic assumption of the genome projects from the late 1980s to the early 2000s was that if one knew the genotype of an individual, one could deduce their phenotype. This was the basic claim underlying their justification to the agencies that funded these projects. In fact, it was already clear that this could not be true.
Take the phenomenon of recessive genes. Everyone has heard that some mutations are recessive, i.e. do not show their effect if there is a normal or “wild-type” allele also present in the genome. Every animal and every plant carries two sets of chromosomes in their body cells, one from the maternal parent, the other set from the paternal side. Hence, each gene is present in two copies. Where the specific forms of a particular gene differ, they are termed “alleles”. Imagine that one form, let us symbolize it as capital B, always expresses its phenotype whether it is accompanied in those body cells by another B allele or another allele, let us denote it as” b”, that produces a different phenotype (when in two copies). If the BB genotype produces an individual that looks the same as one produced by a Bb individual, we say that B is dominant to b and that b is recessive (to B).
The crucial point is that we can never deduce these outcomes from the genotype, even having the whole DNA sequences of the organism or even just the gene in its two copies. What we can say is that most recessive alleles are due to
l-o-f mutations AND that when we see the phenomenon of dominance, it is because one wild-type allele is sufficient to produce the B phenotype. The crucial point is that one cannot predict dominance or recessiveness from the sequence alone, though with a lot more information the sequence provides a clue. More importantly, dominance and recessiveness are not a universal rule. For some genes, the phenotype produced from the Bb genotype is clearly intermediate between those produced by BB and bb respectively. In still other cases, the phenotype produced by Bb has novel features.
I will spare you the jargon name of this 50% gene activity condition but give an example. There is a gene that when present in only one functional copy creates people with certain distinctive facial characteristics AND increased sociability. These individuals love to be around other people, talk readily and often interestingly, though they are often slightly cognitively impaired.6 Hence, the reduction by 50% of this gene’s activity leads to a distinctive new property, increased sociability!
Many of the genes that show a new biological property with only 50% normal activity are termed “regulatory genes”, namely genes that either turn-up or dial-down other genes’ output of their messenger RNAs. In the case of this example, we can infer that the normal full activity of this gene serves to change the underlying neural circuitry involved in social interactions: that full activity damps down the person’s sociability. Evidently, a loss-of-gene activity can promote a new biological property, in this case, sociability.
There are two key points here. First, that one cannot predict the phenotype from the genotype. Second, that the amount of the gene activity is crucial for determining what the phenotype will be. To be sure, there are situations where a single l-o-f gene change can lead unambiguously to a predictable phenotype. Such single-gene conditions are termed Mendelian, after Gregor Mendel (see footnote 4). Disease conditions such as sickle cell anemia, phenylketonuria, cystic fibrosis, and Huntington’s chorea, are in this category, and show straightforward dominance-recessivity rules for their alleles. But, in general, for most traits and for most disease conditions, the phenotype cannot allow you to predict the genotype, just as the reverse is not possible either.
So far, we have only been talking about the effects of changes in gene activity of single genes and already there should be an idea of how diverse the phenotypic changes can be from altered gene activities. Yet, one of the big advances in 20th century genetics was moving from single gene affects to understanding how so many traits in humans (and, likewise, in all animals and plants) are built on the activities of many, many genes, often in the hundreds to thousands. This is the fact of polygenic determination of phenotypes mentioned earlier. An example is human height. It has been shown by sophisticated methods, which we will not get into here, that show that over 500 human genes (out of the 21,000 protein-coding genes found in the human genome) affect human height, as determined by the collective affects of their particular alleles. Each such gene has a small effect, indeed often a very small effect, but put them all together and depending upon which alleles are involved, the effects can be either to produce a very short human being or a very tall one. (There are also a few genes whose alleles can alter height dramatically but most of the variation in human height is cumulative from individually small effects of hundreds of genes.)
That example indicates one form of polygenic effect, namely additive effects. The other kind, far more important in scope, involves interactive effects between genes. The great majority of the protein-coding genes in the human genome (a total of about 21,000) exert their effects in this way, by interacting with and regulating (up or down) the activity of other genes. For the development of a given physical trait, such as a lung or the heart, there are hundreds to thousands of such interacting genes.
These sets of interacting genes are termed “genetic regulatory networks” or GRNs. In principle, changes in gene activity through genetic variants in most of these genes can effect the operation of the whole network and thus affect how the trait develops. Each such network evolves with time in the developmental process that creates the trait – whether it is an organ like the lung or the heart or the neural circuitry that underlies a complex behavior – and any interference with any gene activity in the early stages of development is bound to have an effect on later stages and the final outcome. Think of it as like a complex business plan with defined players who each must do something at a particular time if the products are to be correct and produced in the right number. If something interferes with one of the players carrying out their role, you can have serious “downstream” effects. GRNs are like that.
This is complex material and one can write a book, not just a 3000-word article, about it. I hope, however, that I have intrigued rather than discouraged you. The relevance of this material will become apparent when we explore neurodivergence and its complexities.
If basic molecular biology is largely terra incognita for you, I recommend either James Watson’s “The Molecular Biology of the Gene”, especially the earlier editions, or “Genetics for Dummies”, 4th edn.
Proteinaceous molecules are the workhorses of the cells. They carry out all the chemical transformations, its “metabolism”, that the body needs for health, growth, and its maintenance. Most proteins consist of two or more polypeptide chains though some polypeptide chains, depending on the gene, act on their own.
For a review, see Mattick, J.S. et al. (2023). Long non-coding RNAs: definitions, foundations, challenges and recommendations. Nat. Rev. Mol. Cell Bio. 24: 430-447.
Few scientific fields have a known birth date; their evolution tends to be gradual without a dramatic starting point. Genetics, in contrast, has two. The first is 1866, when Gregor Mendel, a Moravian friar and amateur plant breeder published his discoveries about the heredity of pea plants. Few people paid attention and Mendel died in 1882, without recognition. His work, however, was rediscovered in 1900 and given full credit by the three men who found it. Generally, this is considered the birth date of genetics as a field.
The terms “genotype” and “phenotype” were coined by Wilhelm Johannsen. He also gave us the word “the gene”. See no. 5, “On how the term the “gene” first gained and then lost its meaning”.
This gene’s activity and role were first elucidated in very friendly dogs! See von Holdt, B. (2017). Structural variants in genes associated with human Williams-Beuren syndrome underlie stereotypical hypersociability in domestic dogs. Sci Adv. doi. 10.1126/sciadv.1700398


