Market Cap: $2.1832T 0.55%
Volume(24h): $55.2152B -3.33%
  • Market Cap: $2.1832T 0.55%
  • Volume(24h): $55.2152B -3.33%
  • Fear & Greed Index:
  • Market Cap: $2.1832T 0.55%
Cryptos
Topics
Cryptospedia
News
CryptosTopics
Videos
Top News
Cryptos
Topics
Cryptospedia
News
CryptosTopics
Videos
bitcoin
bitcoin

$87959.907984 USD

1.34%

ethereum
ethereum

$2920.497338 USD

3.04%

tether
tether

$0.999775 USD

0.00%

xrp
xrp

$2.237324 USD

8.12%

bnb
bnb

$860.243768 USD

0.90%

solana
solana

$138.089498 USD

5.43%

usd-coin
usd-coin

$0.999807 USD

0.01%

tron
tron

$0.272801 USD

-1.53%

dogecoin
dogecoin

$0.150904 USD

2.96%

cardano
cardano

$0.421635 USD

1.97%

hyperliquid
hyperliquid

$32.152445 USD

2.23%

bitcoin-cash
bitcoin-cash

$533.301069 USD

-1.94%

chainlink
chainlink

$12.953417 USD

2.68%

unus-sed-leo
unus-sed-leo

$9.535951 USD

0.73%

zcash
zcash

$521.483386 USD

-2.87%

Cryptocurrency News Articles

Andrej Karpathy Believes That LLMs Will Make Deep Learning Frameworks Like PyTorch Redundant

Sep 17, 2024 at 04:04 pm

LLMs are becoming the default solution for many of the problems that businesses and researchers face. Even when it comes to domains outside language and text, people have been experimenting with LLMs for guessing and predicting the next token, which sparks an interesting conversation around the need for other tools like PyTorch, as LLMs can just do the whole job in the future.

Andrej Karpathy Believes That LLMs Will Make Deep Learning Frameworks Like PyTorch Redundant

The rapid advancement of large language models (LLMs) has sparked a discussion about their potential to replace other tools, such as deep learning frameworks like PyTorch, for a wide range of problems.

LLMs are typically designed to predict the next token in a sequence, whether that sequence consists of words, images, or other types of information. This next token prediction framework can be applied to a diverse set of problems, extending beyond text.

Andrej Karpathy suggests that deep learning frameworks, like PyTorch and its counterparts, might be overly general for the majority of problems in the future.

LLMs have evolved beyond their initial specialization in language. The term “language” is used historically because these models were first trained to predict the next word in a sentence. However, LLMs can work on any kind of data that’s broken down into small pieces, called tokens.

Imagine LLMs as a super-smart guessing game. If you’re building a car, a house, or an animal using Legos, you’re just putting blocks together. LLMs don’t care if the tokens (blocks) represent words, images, or even molecules—they just focus on predicting what the next block should be based on what’s already there.

For example, protein prediction models like AlphaFold and ESMFold are built on top of generative language models. Calling such intricate models LLMs might be limiting.

Karpathy further highlights that the term “language” might be leading people to believe that LLMs are limited to text applications, which is not the case.

“I don’t think this is true but I think it’s half true,” Karpathy said in response to a thread about LLMs being able to handle almost all types of problems.

“Probably the name should change.”

“Definitely needs a new name. ‘Multimodal LLM’ is extra silly, as the first word contradicts the third word,” replied Elon Musk in the same thread.

Meanwhile, Yann LeCun is more concerned about why this doesn’t make sense for all the types of problems.

“It only works with discretized outputs (discrete symbols) and only makes sense with symbol sequences with a natural order (not images). Text, DNA, proteins, musical scores, etc. are discrete or easily discretized,” said LeCun.

For something like images, which are continuous and don’t naturally have a strict sequence of discrete symbols (each pixel doesn’t follow a clear ‘order’ like text), LLMs don’t work as naturally. To use an LLM for images, you would first need to somehow convert the image into discrete chunks (like dividing the image into small patches), but this doesn’t follow the same natural order that exists in text or DNA.

Agreeing with Karpathy, and a little with LeCun, Gary Marcus said that statistical modelling of token streams works well if reasoning or planning isn’t required.

Last month, Eliezer Yudkowsky also said that predicting the next token can solve almost all the well-posed problems.

“literally any well-posed problem is isomorphic to ‘predict the next token of the answer’,” he said.

“Throwing an LLM at it”

The idea that many problems can be reduced to a token-stream prediction model is intriguing, especially since domains like images, audio, and even molecules can be broken down into sequences of tokens. This suggests that a unified approach like LLMs could handle diverse tasks, reducing the need for highly specialised architectures, such as PyTorch.

However, frameworks like PyTorch provide more than just flexibility in creating neural network models. They allow for a variety of deep learning operations that aren’t necessarily relevant for LLMs but are critical for other areas like reinforcement learning, generative models, and non-sequential tasks.

While it’s true that LLMs could dominate many applications, not every problem is best framed as “next token prediction.” We may see a simplification or specialization of deep learning frameworks to accommodate the increasing dominance of LLM-based models. Still, the complete redundancy of frameworks like PyTorch might be too extreme of a prediction.

OpenAI’s newest model o1 gives a sense of why LLMs would be able to solve a lot of problems outside of the realm that is currently considered achievable. With the reasoning tokens in place, the model can go beyond just ‘predicting’ the next token and giving reasons for why it did so.

PyTorch and similar frameworks may not become redundant but could evolve to become more focused on token-based models, while still offering tools for more diverse problems outside that paradigm.

Though, currently only in language or text format, the capabilities might extend beyond it soon. Calling LLMs as LLMs might be underrepresenting their capabilities. Moreover, “it just predicts the next token” is a thought-terminating cliche.

Original source:analyticsindiamag

Disclaimer:info@kdj.com

The information provided is not trading advice. kdj.com does not assume any responsibility for any investments made based on the information provided in this article. Cryptocurrencies are highly volatile and it is highly recommended that you invest with caution after thorough research!

If you believe that the content used on this website infringes your copyright, please contact us immediately (info@kdj.com) and we will delete it promptly.

Other articles published on Aug 06, 2026