GPT-3 A Hitchhiker's Guide

GPT-3 A Hitchhiker's Guide

July 20, 2020• 10 min read

The goal of this post is to guide your thinking on GPT-3. This post will:

If you think something should be added, email m@lambdalabs.com.

What researchers are saying about GPT-3

" For me, the big story about #gpt3 is not that it is smart - it is dumb as a pile of rocks - but that piles of rocks can do many things we thought you needed to be smart for. Fake intelligence may be dominant over real intelligence in many domains."

– Anders Sandberg, Senior Research Fellow at Oxford University, read tweet

"_The GPT-3 hype is way too much. It’s impressive (thanks for the nice compliments!) but it still has serious weaknesses and sometimes makes very silly mistakes. AI is going to change the world, but GPT-3 is just a very early glimpse. We have a lot still to figure out. _"

– Sam Altman, CEO of OpenAI, read tweet

" I've stayed away from Twitter waiting for the GPT-3 mist to fade.

Good news: More and more people are excited than ever about the possibilities of language modeling Love letter

Bad news: There's a rush of hot takes that forget the progression of the field and make tea leaves of science."

– Stephen Merity, Former Senior Researcher at SalesForce, read tweet

"The transformer architecture of GPT upper bounds its ability at memorization. It cannot learn many algorithms due to the functional form of its forward pass, and spends a fixed compute per token - i.e. it can't "think for a while". Progress here critical, likely but non-trivial."

– Andrej Karpathy, Senior Director of A.I. at Tesla, read tweet

"I used to say that AI research seemed to have an odd blind spot towards automation of programming work, and I suspected a subconscious self-preservation bias. The recent, almost accidental, discovery that GPT-3 can sort of write code does generate a slight shiver."

– John Carmack, Consulting CTO of Oculus VR, read tweet

"GPT-3 often performs like a clever student who hasn't done their reading trying to bullshit their way through an exam. Some well-known facts, some half-truths, and some straight lies, strung together in what first looks like a smooth narrative."

– Julian Togelius, Associate Professor researching A.I. at NYU, read tweet

"Hackers are fascinated by GPT-3. To everyone else it seems a toy. Pattern seem familiar to anyone?"

– Paul Graham, Founder of YCombinator, read tweet

"Our exciting future of security vulnerabilities involving GPT-3 like models in web apps..."

– Chris Olah, Member of Technical Staff at OpenAI, read full tweet

"Got my invite to the @OpenAI GPT-3 API from @gdb. I actually think it deserves more hype than it’s getting, but not necessarily for the magical reasons Twitter touts. Why? My quick thoughts and impressions: (1/11)"

– Shreya Shankar, ML researcher at Viaduct AI, read full tweet.

The best GPT-3 technical write-ups

OpenAI's GPT-3 Language Model: A Technical Overview

June 3, 2020 by Chuan Li - Chief Science Officer at Lambda

Link

An overview of the original paper covering its training cost and research implications.


GPT-3

July 18, 2020, by Gwern Branwen

Link

A well-cited overview of the original GPT-3 paper with a punchline on what it tells us about the scaling hypothesis:


Tempering Expectations for GPT-3 and OpenAI’s API

July 18, 2020, by Max Woolf - Data Scientist at BuzzFeed, Ex- Apple

Link

A hacker's introduction to GPT-3 with a curl-based examples of OpenAI's invite-only beta API:


Why GPT-3 Matters

Link


GPT-3: A Disappointing Paper?

Link

How GPT-3 Works

Link

A visual introduction to GPT-3.

GPT-3: Language Models are Few-Shot Learners

Link

The original GPT-3 paper from OpenAI.

Other interesting reads

The best videos on GPT-3

* GPT-3: Language Models are Few-Shot Learners (Paper Explained)

Watch video

OpenAI GPT-3: Language Models are Few-Shot Learners

Watch video

GPT 3 Demo and Explanation - An AI revolution from OpenAI

Watch video

GPT3: An Even Bigger Language Model

Watch video

Cool GPT-3 demos