r/MachineLearning
r/MachineLearning

r/MachineLearning

r/MachineLearning is a Reddit community for machine learning researchers and enthusiasts. It features discussions on topics like networking at conferences and the long-term value of AI research.

Profile generated by AI for Anything

Loss functions in Instance Representation Learning [R]

Loss functions in Instance Representation Learning [R]

13d

I'm trying to implement CALM paper, and I have some questions. [P]

I'm trying to implement CALM paper, and I have some questions. [P]

13d

The original title is about a quiz that tells you which LLM you align with most. Let me rewrite this for a mobile feed.

The original title is about a quiz that tells you which LLM you align with most. Let me rewrite this for a mobile feed.

14d

NagaTranslate: Building a translation and voice pipeline for low-resource Nagaland creoles (Whisper, VITS, LLMs) [P]

NagaTranslate: Building a translation and voice pipeline for low-resource Nagaland creoles (Whisper, VITS, LLMs) [P]

15d

Attention pathologies stem from norm

Attention pathologies stem from norm

17d

[R] Compiling Agentic Workflows into LLM Weights: Near-Frontier Quality at Two Orders of Magnitude Less Cost

[R] Compiling Agentic Workflows into LLM Weights: Near-Frontier Quality at Two Orders of Magnitude Less Cost

17d

[R] All Routes Lead to Collapse: attention sinks, representation collapse, and norm stratification are what content-based routing does under a norm-blind metric

[R] All Routes Lead to Collapse: attention sinks, representation collapse, and norm stratification are what content-based routing does under a norm-blind metric

17d

CALHippo - Mapping neurons and glial cells in the human brain hippocampus in 3D using SOTA segmentation and density estimation models [R]

CALHippo - Mapping neurons and glial cells in the human brain hippocampus in 3D using SOTA segmentation and density estimation models [R]

18d

High Dimensional, Dynamic Rotary Positional Embedding [P]

High Dimensional, Dynamic Rotary Positional Embedding [P]

18d

I compiled LLM inference pricing across 7 providers — the caching numbers are surprising(spreadsheet included) [R]

I compiled LLM inference pricing across 7 providers — the caching numbers are surprising(spreadsheet included) [R]

19d

DeepSWE: new benchmark looking at how well today's frontier models can actually write code [R]

DeepSWE: new benchmark looking at how well today's frontier models can actually write code [R]

19d

Free lecture covers multivariate probability

Free lecture covers multivariate probability

25d

Next-Latent Prediction Transformers [R]

Next-Latent Prediction Transformers [R]

26d

What is Speculative Decoding? (trending on paperswithco.de) [R]

What is Speculative Decoding? (trending on paperswithco.de) [R]

26d

AI language models have favorite names, and we mapped them [R]

AI language models have favorite names, and we mapped them [R]

27d

go from zero to claude code pro in one day — hands on bootcamp may 30 [D]

go from zero to claude code pro in one day — hands on bootcamp may 30 [D]

63d

Signals: finding the most informative agent traces without LLM judges [R]

Signals: finding the most informative agent traces without LLM judges [R]

63d

Steam Simularity Reccomender Student Project [p]

Steam Simularity Reccomender Student Project [p]

64d

Steam Similarity Recommender Find your next favorite game and learn WHY (student project)[P]

Steam Similarity Recommender Find your next favorite game and learn WHY (student project)[P]

65d

Formalizing statistical learning theory in Lean 4 [R]

Formalizing statistical learning theory in Lean 4 [R]

65d

What should a PyTorch training end-of-run performance summary show? [D]

What should a PyTorch training end-of-run performance summary show? [D]

66d

Steam Similarity Recommender [P]

Steam Similarity Recommender [P]

66d

Transformer Math Explorer [P]

Transformer Math Explorer [P]

66d

Question about PLS-DA hyperparameter tuning [R]

Question about PLS-DA hyperparameter tuning [R]

68d