# Modern Artificial Intelligence in the 1970s: The Past Was Already Thinking
![[resources/images/modern-ai-in-the-70s.png]]
[[wiki/Lawrence Roberts|Lawrence Roberts]] wanted ten thousand words, spoken by anyone. That was the original ambition when [[wiki/DARPA|ARPA]] initiated its **[[wiki/Speech Understanding Research|Speech Understanding Research]]** program in 1971. The advisory board he convened — chaired by [[wiki/Allen Newell|Allen Newell]], with [[wiki/J. C. R. Licklider|J. C. R. Licklider]] among its members — reduced the objective to something a committee could defend: **a thousand words, in a quiet room, by a limited number of speakers, within a restricted subject vocabulary**. Roberts committed three million dollars a year for five years and intended to fund a second five-year program after it.
Five years later, Carnegie Mellon's **[[wiki/HARPY|HARPY]]** was operating on **connected speech across a 1,011-word vocabulary**, tested across multiple speakers, running on a PDP-10-class machine. The surviving CMU material sets the 1971 targets directly against the 1976 measurements and states the result without hedging: the system "not only satisfies the original goals, but exceeds some of the stated objectives." The National Academies' history of federally funded computing research concurs that HARPY came closest to the original goals and may have exceeded the program's benchmarks. Controversy then arose over how the testing had been conducted, some [[wiki/DARPA|DARPA]] program managers concluded that SUR had failed, and **the intended follow-on was never funded**. Licklider's own retrospective assessment survives in a footnote of that history: the speech project met its main objectives, but that was not enough to save it.
Every clause of that sequence is documentary rather than interpretive, and the sequence is the argument. **A target was negotiated downward by committee. A machine reached the headline benchmark and possibly exceeded the program's objectives. The principal architect of the initiative said afterward that its main objectives had been met. It was terminated regardless.** Capability had been demonstrated. Capability was not sufficient.
That is the first fracture in the received chronology, and it does not close.
## What the termination built
It would be too easy — and too weak — to read HARPY as a story about a good system losing its funding. The more consequential fact is what the dispute produced downstream.
The controversy was not principally about whether the machine worked. It was about **whether the way the machine had been tested licensed the claim that it worked**. Vocabulary size was only one of the program's dimensions; speaker breadth, semantic error rate, processing speed, syntactic and task constraints, and the degree of per-speaker tuning all bore on the verdict, and reasonable participants disagreed about how to weigh them. Capability did enter the adjudication. What the record establishes is narrower and far more damaging: **demonstrated capability, and even attainment of the program's stated main objectives, were not sufficient to determine continuation.**
The National Academies traces a direct line from that episode to a later DARPA regime in which evaluation methods were **prescribed in advance**, with annual system evaluations against fixed benchmark tasks — and notes that critics regarded this as part of a broader migration away from open-ended basic research toward narrowly measurable objectives. Which means the entire evaluative apparatus that now governs [[wiki/Machine Learning|machine learning]] — the [[wiki/AI Benchmarking|shared benchmark, the leaderboard, the annual bake-off, the metric agreed upon before the work begins]] — descends in part from an argument about whether a successful 1970s system had actually succeeded.
Standardized testing did not begin here, and the claim should not be inflated into one. But **one of the decisive institutional lineages of modern benchmark culture runs directly out of a dispute over whether a machine that met its target had met its target.** That is not a footnote to the history of artificial intelligence. It is an origin point for the administrative epistemology under which the field has operated ever since, and it sits inside the decade that popular memory files under failure.
## The vocabulary problem
Now open [[wiki/Geoffrey Hinton|Geoffrey Hinton's]] doctoral thesis, _Relaxation and Its Role in Vision_, accepted at the University of Edinburgh in **1977** under Christopher Longuet-Higgins. On printed page 69, having constructed a particular space of feasible states, Hinton observes that [[wiki/Local Optimum|local maxima]] do not arise within it, and concludes that "the usual objection to [[wiki/Hill Climbing|hill-climbing]], that it gets stuck at local maxima" therefore does not apply to his procedure.
Read that with the date removed. A researcher is reasoning about the **topology of an [[wiki/Optimization Landscape|optimization landscape]]**, identifying the **local-optimum pathology** as the standing objection to gradient-following search, and arguing that the geometry of his particular construction exempts him from it. The scholarly restraint required is minimal but real: Hinton is proving the absence of local maxima **in that space**, not announcing a general property of learning landscapes. What the sentence establishes beyond dispute is that in 1977 the local-minimum problem was not a future concern awaiting the neural-network revival. It was a known objection that a doctoral candidate had to answer — and, as the next section shows, one that had been under discussion in the published literature for a decade before he answered it.
The thesis abstract states the problem in terms that require almost no translation: selecting the best consistent combination from among many interrelated, locally plausible hypotheses about how parts of a visual input may be interpreted, with each hypothesis assigned a **supposition value between 0 and 1** and a **[[wiki/Relaxation Labeling|parallel relaxation operator]]** grounded in plausibilities and logical relations iteratively modifying those values until the surviving consistent set approaches unity and the remainder collapses toward zero. The preceding year, in _Using Relaxation to Find a Puppet_, the ostensible task had been recognizing a puppet in configurations of overlapping rectangles; the actual task was **evading [[wiki/Combinatorial Explosion|combinatorial enumeration]] by letting local interpretations negotiate in parallel toward global coherence**.
The habitual description of this work — early research that would later influence neural networks — performs a quiet substitution. It converts a researcher **doing** machine intelligence into a researcher **anticipating** it.
## What was actually running
The substitution becomes indefensible once the thesis is restored to its neighborhood — and the first thing that neighborhood does is leak across the boundary this essay's title depends on.
**The decade is a container, not a fact.** [[wiki/Alexey Ivakhnenko|Ivakhnenko]] and Lapa's _Cybernetic Predicting Devices_ appeared in **1965**. [[wiki/Shun-Ichi Amari|Amari's]] adaptive pattern classifier appeared in **1967**. [[wiki/Kunihiko Fukushima|Fukushima's]] multilayered network of analog threshold elements, credited by Schmidhuber with the first use of what is now called the [[wiki/Rectified Linear Unit|rectified linear unit]], appeared in **1969**. [[wiki/Marvin Minsky|Minsky]] and [[wiki/Seymour Papert|Papert's]] _Perceptrons_, the document usually blamed for the [[wiki/Connectionism|connectionist]] recession, appeared in **1969**. The first [[wiki/ARPANET|ARPANET]] message was transmitted in October of **1969**. [[wiki/Seppo Linnainmaa|Linnainmaa's]] reverse-mode thesis arrives in **1970** — three months on the far side of a line drawn by the base-ten notation of the Gregorian calendar and nothing else. The 1970s are a convenient frame for an argument, and the frame does not survive contact with the material it contains, which is itself a small demonstration of the thesis: **[[wiki/Periodization|periodization]] is a narrative technology, and it is never neutral.**
But the frame is not arbitrary either, and the reason is worth stating precisely. The 1960s supply **antecedents**; the early 1970s supply **simultaneity**. In 1972 alone, **four papers on correlation-type associative memory appeared from three continents**: Anderson in the United States, Kohonen in Finland in April, Nakano in Tokyo roughly three months later, and Amari, also in Tokyo, in the same year. Amari's own retrospective account of the period names exactly that quartet. A field is not identifiable by its first instance. It is identifiable by the moment **separate groups on separate continents arrive at the same architecture inside the same twelve months**, which is the signature of a shared problem being worked rather than an idea being had. That signature appears at the opening of the decade and does not subsequently disappear.
There is a second observation the boundary makes available. _Perceptrons_ is conventionally assigned responsibility for suppressing connectionist research after 1969, and the years immediately following it are the densest in this entire account. The tidy explanation is that a book extinguished a research program. Whatever role _Perceptrons_ actually played in the contraction of connectionist funding and prestige in the United States and Britain — and historians of the episode dispute the degree considerably — it plainly did not terminate the research program globally, because Soviet and Japanese laboratories operating outside that discourse went on building, and the evidence is Ivakhnenko, Amari, Nakano and Fukushima. The received story is not so much false as **provincial**: an account of one funding ecology mistaken for an account of a global science.
In **1971**, Alexey Ivakhnenko published _Polynomial Theory of Complex Systems_ in _IEEE Transactions on Systems, Man, and Cybernetics_: a multilayered network structure whose successive layers compose progressively more complex combinations of input variables to approximate a decision hypersurface, whose parameters are estimated from data, and whose candidate elements are retained or discarded according to performance on a **[[wiki/Held-Out Validation|separate testing set]]** — a procedure he named _decision regularization_. Held-out validation, deployed as a defense against [[wiki/Overfitting|overfitting]], inside a learned multilayer architecture, published in 1971. The underlying **[[wiki/Group Method of Data Handling|Group Method of Data Handling]]** had been described the previous year as heuristic self-organization.
Shun-Ichi Amari's 1967 _A Theory of Adaptive Pattern Classifiers_ develops adaptive weight learning under a steepest-descent rule whose stochastic approximation he calls a **probabilistic descent method**. Schmidhuber's reading — that this constitutes the first [[wiki/Stochastic Gradient Descent|stochastic gradient descent]] procedure applicable to training multilayer networks end to end — is an interpretation worth attributing to him rather than asserting flatly, since the paper's own text does not deploy the multilayer vocabulary. But the primary document contains something arguably more useful to the present argument: Amari explicitly reasons about the assumption that the average risk possesses **no non-global local minima**. The optimization-landscape question, which the viral version of this history places in the 1980s and the careful version places in Hinton's 1977 thesis, is being handled in a published paper in **1967**.
By **1972**, Amari had published learning [[wiki/Recurrent Neural Network|recurrent networks]]. The same year, [[wiki/James A. Anderson|James Anderson's]] _A Simple Neural Network Generating an Interactive Memory_ modified connection strengths according to the joint activity of connected units so the system could reconstruct an associated pattern from a partial or corrupted cue, with the noise behavior mathematically analyzed; and [[wiki/Teuvo Kohonen|Teuvo Kohonen's]] _Correlation Matrix Memories_ distributed stored information across a matrix rather than depositing it at addresses. [[wiki/Kaoru Nakano|Kaoru Nakano's]] _Associatron_ stored entities as distributed bit patterns and recalled the whole from a fragment, degrading gracefully as load increased; Nakano had already reported simulating a 180-neuron version and constructing a **25-neuron hardware realization**, and framed the entire exercise explicitly as association mechanisms usable in machine intelligence. Amari has since observed that the model he published that same year was the same as what would later be known as the **[[wiki/Hopfield Network|Hopfield model]]** of [[wiki/Associative Memory|associative memory]] — a claim of his own about his own work, made in 1972 and recognized in 1982. **[[wiki/Distributed Representation|Distributed representation]], associative recall, and learning-as-weight-modification** were instrumented, published, built in hardware and analyzed while the Nixon administration was in office.
The backward-derivative lineage requires more care than it usually receives, because it is one of the most contested priority disputes in the field and the popular account botches it in both directions. Seppo Linnainmaa's 1970 master's thesis at Helsinki gives an early general formulation and implementation of **[[wiki/Reverse-Mode Automatic Differentiation|reverse-mode automatic differentiation]]** — with FORTRAN code — published in English in _BIT_ in 1976. [[wiki/Paul Werbos|Paul Werbos's]] 1974 Harvard dissertation, _Beyond Regression_, independently develops what he calls **dynamic feedback**, computing all derivatives in a single pass by proceeding "backwards down any ordered table of operations," provided those operations are differentiable. That is reverse accumulation through composed nonlinear computation, and any account claiming Werbos lacked an efficient method in 1974 has overcorrected into error.
The genuinely open question is narrower: when the machinery received its **explicit worked application to training an artificial neural network**. Werbos himself files his 1981 conference contribution — published in 1982 as _Applications of Advances in Nonlinear Sensitivity Analysis_ — under the heading of backwards differentiation used in training neural network systems, and Schmidhuber's chronology dates the first neural-network application there. That dating belongs to Schmidhuber and should be attributed to him rather than issued as settled historiography. Andreas Griewank's history of reverse-mode differentiation refuses the single-inventor structure entirely, describing **multiple independent incarnations**, with work by Ostrowski and collaborators preceding Linnainmaa in a different context.
Which leaves the fact that actually matters, and it is stronger than any priority claim: **by the middle of the 1970s, at least two independent intellectual lineages, working from unrelated problem domains, had converged on efficient backward credit assignment through composed nonlinear computation.** Convergence of that kind is not what a field looks like before it exists.
In **1975**, Kunihiko Fukushima built and ran the **[[wiki/Cognitron|Cognitron]]**, a self-organizing multilayered neural network simulated on a digital computer, in which connection strengths changed through repeated exposure and deeper layers developed progressively broader receptive fields responding selectively to learned patterns — with no teacher specifying any unit's desired output. [[wiki/Stephen Grossberg|Stephen Grossberg's]] 1976 papers on **adaptive pattern classification and universal recoding** developed the mathematics of feature-detector formation, feedback and expectation in parallel adaptive systems. And in August **1979**, at the Sixth International Joint Conference on Artificial Intelligence in Tokyo, Fukushima presented an architecture of alternating plastic **S-cells** and pooling **C-cells** that self-organized without supervision into representations largely invariant to the **position** of the stimulus. He called it the **[[wiki/Neocognitron|neocognitron]]**. It is the direct ancestor of the [[wiki/Convolutional Neural Network|convolutional network]], and it was demonstrated at a general artificial-intelligence conference before the decade was out.
None of this is a catalogue of predictions. These are **compiled programs, learning rules, convergence analyses, held-out evaluation protocols and measured invariances**.
## Credit assignment, applied to the historians
Honesty about provenance is compulsory in an essay whose subject is misattributed provenance. The Ivakhnenko–Amari–Linnainmaa–Werbos–Fukushima sequence is not my discovery. It is the standing revisionist account assembled over more than a decade by **[[wiki/Jürgen Schmidhuber|Jürgen Schmidhuber]]**, most completely in _Annotated History of Modern AI and Deep Learning_, which opens by declaring that machine learning is the science of [[wiki/Credit Assignment|credit assignment]] and immediately extends the problem to historians interpreting the present through prior contributions. In December 2024 he issued Technical Report IDSIA-24-24, _A Nobel Prize for Plagiarism_, contesting the credit structure of that year's physics award — an allegation and an argument, not an adjudicated finding, and it should be read as his.
One need not adopt his verdicts to register the recursion, which is the quietest and most devastating fact in this territory. **The discipline whose central technical problem is the assignment of credit backward through a deep composition of prior contributions has proven unable to perform that operation on its own history.** The gradient does not flow. The early layers receive nothing. And Griewank's refusal to name a single inventor of reverse-mode differentiation is the methodologically correct posture generalized: in a sufficiently active field, **credit is distributed rather than located**, and any history that resolves to a single name has almost certainly lost information.
## How a decade disappears
Two mechanisms erase this period. Only one of them is discussed.
The familiar mechanism is semantic. Larry Tesler's formulation — intelligence is whatever machines have not done yet — and Pamela McCorduck's observation that each success is promptly reclassified as mere computation together constitute the **[[wiki/AI Effect|AI effect]]**. Associative memory becomes linear algebra. Machine perception becomes image processing. Reverse-mode differentiation becomes numerical analysis. Self-organization becomes [[wiki/Pattern Recognition Systems|pattern recognition]]. Planning becomes search. Speech becomes signal processing. Every maturing capability graduates into a narrower and more respectable discipline, and the vacated category is refilled with whatever remains undone. **A field defined by its residue is perpetually new by construction.**
The unfamiliar mechanism is administrative, and HARPY exposes it. Capability in this period was not yet settled by the kind of standardized, pre-specified evaluation regime that the SUR controversy would help institutionalize; the full testing protocol had not been fixed at the outset, which is precisely what made attainment arguable after the fact. **The regime was manufactured by the dispute rather than available to settle it.** Judgment was rendered by committee. The [[wiki/Lighthill Report|Lighthill report]] of 1973, commissioned by Britain's Science Research Council, partitioned the field into advanced automation, central-nervous-system research and the robot-building work intended to bridge them, and its skepticism about combinatorial explosion in the bridging category **helped withdraw the institutional floor beneath** British artificial intelligence for the better part of a decade. Edinburgh's own institutional history describes a loss of confidence lasting roughly that long; more recent historical work complicates the direct causal attribution without disturbing the outcome. Three years later the American speech program was terminated with its headline objective met.
The 1970s did not fail. **The 1970s were adjudicated** — by funding bodies, review panels and disciplinary politics — and the adjudication was subsequently laundered into a technical verdict, remembered as a season rather than a decision. The word _[[wiki/AI Winter|winter]]_ does considerable work here, because winters are meteorological. Nobody is responsible for a winter.
## The aeronautical objection, and why it self-destructs
The reflexive rebuttal is that none of this operated at consequential scale and therefore none of it was really machine intelligence. The objection cannot be stated in a form that survives its own specification.
Name the year aviation began. Say **1903** and you have accepted that **twelve seconds and a hundred and twenty feet at under seven miles an hour** constitutes flight — at which point consequential capability has been abandoned as the criterion and demonstrated mechanism has replaced it, and the Cognitron qualifies on identical grounds. Say 1914, or 1939, or the jet age, and you are dating a physical science by **passenger comfort and market penetration**, which is precisely the visibility-indexed historiography under examination. Say 1799, when George Cayley separated lift from thrust and drag, or the 1890s, when Otto Lilienthal accumulated some two thousand glides, and you have conceded the principle without argument. There is no fourth answer.
The distinction that carries weight is between **speculative morphology** and **demonstrated mechanism**. Leonardo's ornithopter is a drawing of an object that could not work, produced without a correct theory of lift; that is a precursor in the honest sense of the word. Lilienthal suspended beneath a cambered wing is not a precursor to aviation. He is **flying**, at the only scale then available, exploiting a physical principle he understood and could measure. Fukushima did not sketch a self-organizing network. He **compiled and ran one**, and measured what its deeper layers had acquired. The wing generated lift, and the instruments recorded it.
## Where the analogy honestly breaks
One asymmetry must be stated plainly, because suppressing it would repeat the offense this essay exists to name.
Aeronautics scaled roughly as its practitioners expected: more power, more span, more lift, continuous with the theory. Machine intelligence did not. It would be false to claim that researchers in the 1970s failed to understand that more computation, more memory, more units and more training material could improve performance; they understood that perfectly well, and said so. What they did not possess was the **[[wiki/Neural Scaling Laws|modern empirical scaling regime]]** — the experimentally established and still incompletely explained finding that comparatively stable architectures, trained across orders of magnitude more compute, data and parameters, exhibit surprisingly regular scaling behavior and eventually display capabilities not evident at small scale — a description of what is observed, which does not require committing to any particular account of whether those transitions are genuine discontinuities or artifacts of how they are measured.
The distinction improves the aeronautical comparison rather than damaging it. **They knew that larger aircraft could fly farther. They did not possess the mapped flight envelope.** The defensible form of the thesis is accordingly that the 1970s possessed the physics and not the envelope — which is exactly the aeronautical situation in 1901: the theory substantially correct, the mechanism demonstrated, the scale absent, and the public entirely certain the whole enterprise was a delusion of cranks.
## What actually changed
Machine intelligence spent five decades beneath interfaces where its operation could always be renamed. It ranked, classified, filtered, routed, forecast, recognized, guided and targeted, and each of those verbs offered semantic shelter. Then a system appeared in an empty text box and occupied **language** — the medium through which human beings most reliably detect the presence of another mind. The shelter collapsed, and hundreds of millions of people encountered computational intelligence as an apparent interlocutor rather than as infrastructure.
That was a real discontinuity: in **scale**, in **usability**, and above all in **visibility**. It was not the beginning of machine intelligence. It was the beginning of machine intelligence as a **mass public encounter**, and the persistent conflation of the two is not an innocent simplification. It licenses the belief that this technology is young, that its governance can therefore wait, and that no one has yet had sufficient time to establish a durable position of advantage inside it.
## The timeline you were issued
What most people carry is not a history. It is a **timeline** — a compressed, ordered, teachable sequence, assembled from the historical record by institutions that had positions to protect while they were assembling it. This requires no conspiracy and no coordination, which is exactly why it is so durable. **Every system that persists performs continuous internal risk assessment and positions itself accordingly** — an agency, a discipline, a corporation, a nation, a cell. What such a system publishes about its own past is one of the instruments of that positioning, and the selection pressure operates whether or not anyone involved intends deception.
The material above is the demonstration rather than the allegation. Lighthill wrote a defensible document. DARPA's move toward prescribed evaluation was a reasonable institutional response to a genuine measurement dispute. Every actor in the HARPY episode can be granted good faith and the aggregate result is unchanged: **a chronology that flatters the present, dates the field from the moment of public visibility, and quietly relocates a decade of working systems into the category of prehistory.** No one wrote that story. It precipitated out of a thousand local decisions about what to fund, what to recognize, what to name, and what to teach — each of them rational for the party making it.
The operational consequence is unsentimental. **A summary is cheaper than a record, and you will be issued the summary by default.** Anyone content to inherit the compressed version is navigating by a map drawn by parties with their own interests in where the roads appear to go, and will consistently misjudge how mature a technology is, how long its advantages have been accumulating, and who has been standing inside it. Understanding reality is expensive, effortful and adversarial to convenience, and it is the only available defense.
This essay corrects a few of the things you may have carried forward. There are certainly others.
The archive says otherwise. The capability existed. **Public authorization lagged.** Institutions decided what would be recognized, what would be continued, what would be classified as success and what would subsequently be narrated as failure — and a public taught to date a field from its own first glimpse of it will systematically underestimate how long the people who were already looking have had to arrange the room.
## Related reading
If you enjoyed this article, you may also enjoy [[articles/The Evolutionary Roots of Silicon Valley|The Evolutionary Roots of Silicon Valley]] and [[articles/Digital Darwinism and the Invisible World of Machine Evolution|Digital Darwinism and the Invisible World of Machine Evolution]].
---
[[about/About Bryant McGill|Bryant McGill]] is a Wall Street Journal and USA Today best-selling author, founder of Simple Reminders, and architect of the Polyphonic Cognitive Ecosystem. He is a Congressionally Recognized Ambassador of Goodwill and a United Nations appointed Global Champion, whose work spans naval intelligence systems, computational linguistics, and civilizational governance architecture.
---
## References
Amari, Shun-Ichi. "[A Theory of Adaptive Pattern Classifiers](https://doi.org/10.1109/PGEC.1967.264666)." _IEEE Transactions on Electronic Computers_ EC-16, no. 3 (1967): 299–307. [Full text](https://people.idsia.ch/~juergen/amari1967.pdf).
Amari, Shun-Ichi. "[Learning Patterns and Pattern Sequences by Self-Organizing Nets of Threshold Elements](https://doi.org/10.1109/T-C.1972.223477)." _IEEE Transactions on Computers_ C-21, no. 11 (1972): 1197–1206.
Anderson, James A. "[A Simple Neural Network Generating an Interactive Memory](https://doi.org/10.1016/0025-5564%2872%2990075-2)." _Mathematical Biosciences_ 14, nos. 3–4 (1972): 197–220.
Fukushima, Kunihiko. "[Cognitron: A Self-Organizing Multilayered Neural Network](https://doi.org/10.1007/BF00342633)." _Biological Cybernetics_ 20, nos. 3–4 (1975): 121–136.
Fukushima, Kunihiko. "Self-Organization of a Neural Network Which Gives Position-Invariant Response." _Proceedings of the Sixth International Joint Conference on Artificial Intelligence_, Tokyo, August 1979: 291–293. Expanded as "[Neocognitron: A Self-Organizing Neural Network Model for a Mechanism of Pattern Recognition Unaffected by Shift in Position](https://doi.org/10.1007/BF00344251)," _Biological Cybernetics_ 36, no. 4 (1980): 193–202.
Griewank, Andreas. "[Who Invented the Reverse Mode of Differentiation?](https://citeseerx.ist.psu.edu/document?doi=8ff0c546aff84566635f0a9a2e01feb3d6588c1c&repid=rep1&type=pdf)" In _Optimization Stories_, _Documenta Mathematica_, Extra Volume ISMP (2012): 389–400.
Grossberg, Stephen. "[Adaptive Pattern Classification and Universal Recoding: I. Parallel Development and Coding of Neural Feature Detectors](https://doi.org/10.1007/BF00344744)." _Biological Cybernetics_ 23, no. 3 (1976): 121–134; and "[II. Feedback, Expectation, Olfaction, Illusions](https://doi.org/10.1007/BF00340335)," _Biological Cybernetics_ 23, no. 4 (1976): 187–202.
Hinton, Geoffrey E. "Using Relaxation to Find a Puppet." _Proceedings of the AISB Summer Conference_ (1976): 148–157. A DOI, 10.5555/3015508.3015523, was assigned later to the digitized proceedings record.
Hinton, Geoffrey E. "[Relaxation and Its Role in Vision](https://era.ed.ac.uk/handle/1842/8121)." PhD thesis, University of Edinburgh, 1977, p. 69. Edinburgh Research Archive, hdl:1842/8121.
Ivakhnenko, A. G. "[Heuristic Self-Organization in Problems of Engineering Cybernetics](https://doi.org/10.1016/0005-1098%2870%2990092-0)." _Automatica_ 6, no. 2 (1970): 207–219.
Ivakhnenko, A. G. "[Polynomial Theory of Complex Systems](https://doi.org/10.1109/TSMC.1971.4308320)." _IEEE Transactions on Systems, Man, and Cybernetics_ SMC-1, no. 4 (1971): 364–378.
Kohonen, Teuvo. "[Correlation Matrix Memories](https://doi.org/10.1109/TC.1972.5008975)." _IEEE Transactions on Computers_ C-21, no. 4 (1972): 353–359.
Lighthill, James. "Artificial Intelligence: A General Survey." In _Artificial Intelligence: A Paper Symposium_. London: Science Research Council, 1973.
Linnainmaa, Seppo. _The Representation of the Cumulative Rounding Error of an Algorithm as a Taylor Expansion of the Local Rounding Errors_. Master's thesis, University of Helsinki, 1970; developed in English as "[Taylor Expansion of the Accumulated Rounding Error](https://doi.org/10.1007/BF01931367)," _BIT_ 16, no. 2 (1976): 146–160.
Lowerre, Bruce, and Raj Reddy. "[The HARPY Speech Understanding System](https://iiif.library.cmu.edu/file/Newell_box00092_fld06359_doc0001/Newell_box00092_fld06359_doc0001.pdf)." Carnegie-Mellon University, 1976. Allen Newell papers, CMU Libraries.
National Research Council. "[Funding a Revolution: Government Support for Computing Research](https://nap.nationalacademies.org/read/6323/chapter/11)," chapter on developments in artificial intelligence. Washington, DC: National Academies Press, 1999.
Nakano, Kaoru. "[Associatron — A Model of Associative Memory](https://doi.org/10.1109/TSMC.1972.4309133)." _IEEE Transactions on Systems, Man, and Cybernetics_ SMC-2, no. 3 (1972): 380–388.
Newell, Allen, et al. _[Speech-Understanding Systems: Final Report of a Study Group](https://rr.cs.cmu.edu/SUR.pdf)_. Advisory report to ARPA, 1971; published Amsterdam: North-Holland, 1973.
Schmidhuber, Jürgen. "[Annotated History of Modern AI and Deep Learning](https://arxiv.org/abs/2212.11279)." Technical Report IDSIA-22-22, arXiv:2212.11279, 2022, revised 2025.
Schmidhuber, Jürgen. "[A Nobel Prize for Plagiarism](https://people.idsia.ch/~juergen/physics-nobel-2024-plagiarism.html)." Technical Report IDSIA-24-24, 7 December 2024.
Werbos, Paul J. _[Beyond Regression: New Tools for Prediction and Analysis in the Behavioral Sciences](https://gwern.net/doc/ai/nn/1974-werbos.pdf)_. PhD dissertation, Harvard University, 1974, p. II-34; and "Applications of Advances in Nonlinear Sensitivity Analysis," in _System Modeling and Optimization: Proceedings of the IFIP Conference_, Springer, 1982. See also his [annotated list of selected papers](https://werbos.com/Neural/Selected_Papers.htm).