r/datascience 1d ago

Weekly Entering & Transitioning - Thread 07 Sep, 2026 - 14 Sep, 2026

Upvotes

Welcome to this week's entering & transitioning thread! This thread is for any questions about getting started, studying, or transitioning into the data science field. Topics include:

  • Learning resources (e.g. books, tutorials, videos)
  • Traditional education (e.g. schools, degrees, electives)
  • Alternative education (e.g. online courses, bootcamps)
  • Job search questions (e.g. resumes, applying, career prospects)
  • Elementary questions (e.g. where to start, what next)

While you wait for answers from the community, check out the FAQ and Resources pages on our wiki. You can also search for answers in past weekly threads.


r/datascience 16h ago

Tools Can anyone suggest a comprehensive intro to LangGraph?

Upvotes

I need to get up to speed on LangGraph within a couple of days. Not enough to say I'm an expert in it or be able to flub that I've shipped on it, but just enough to be able to competently say I understand the concepts and how it works. Maybe be able to elaborate on how I would have used it in the past (my last gig was AI development intensive but was all hand rolled on the Anthropic API).

I was hoping some folks here could direct me to some resources where I could invest three hours or so and get a nice deep dive on the subject. Youtube videos, tutorials, that sort of thing.

TIA!


r/datascience 1d ago

Discussion Lossless compression that stays queryable (for tables/logs, not JPEG)

Upvotes

Sharing a tool I’ve been watching for data eng / analytics folks: TinyFiles (AT-1) does lossless compression that’s still queryable and verifiable — aimed at archives you actually need to touch again, not media codecs.

Free signup, no card, no sales call: https://tinyfiles.io

Honest caveat: not for JPEG/video.

If you’ve ever compressed a dataset and then couldn’t cheaply query it without unpacking the whole thing, this is the niche.


r/datascience 2d ago

Discussion How do you stay up to date with the latest and greatest?

Upvotes

Aside from the TLDR Newsletter, DevNavigator Newsletter, and various forums like OpenAI and Anthropic's blogs, how is everyone else staying up to date with whats happening in Data Science and AI?


r/datascience 3d ago

Discussion Artificial deadlines part 1, evidence of fraud in an influential study about procrastination

Thumbnail
datacolada.org
Upvotes

r/datascience 3d ago

Discussion Cost-optimal design under heterogeneous treatment cost

Thumbnail
towardsdatascience.com
Upvotes

I wrote a piece on cost optimal designs when treating costs more (literally dollars) than control. I think this topic is too little discussed in mainstream data science, and may be useful for you here. Curious to hear your thoughts on this. any feedback on the article is welcome too.


r/datascience 4d ago

ML How are LLMs used in predictive modeling and anomaly detection?

Upvotes

I am DS with 20 YOE but I've been managing lately and know very little about using LLMs. Does the DS just feed the data into LLM then ask it to predict something or find anomaly? Or does the DS ask the LLM to build a model which is then deployed? Thanks.


r/datascience 6d ago

Discussion At senior levels, where do you draw the line between Data Science, Data Engineering, and Platform ownership?

Upvotes

TL;DR: My DS/analytics role has expanded into senior-level data/platform engineering and client leadership, but my title, pay, and promotion path haven’t kept up.

Update: Had a really good 1:1 today. My boss agreed my scope has grown beyond the JD, asked me to document the differences, and wants me to draft a senior-level JD that reflects what I’m actually doing. We also discussed hiring someone junior under me to take on some analytics work. He wants to use this to support a strong year-end rating and promotion, though he says off-cycle promotions are currently blocked by policy.

I am going to keep applying externally though. Nothing in writing yet.

I’ve been in data science/analytics for about 11 years. Most of my earlier career was at a Fortune 100 financial company, where I eventually became a Data Science Manager and led a small team forecasting risk metrics that fed into public earnings reporting. I’m now a Big Data Analytics Manager at a fintech/fraud prevention company, working remotely in the US.

The reason I’m posting is that my job has changed pretty dramatically from what I was hired to do, and I’m having trouble figuring out what the role actually is anymore. My official job description is basically an Implementation Manager description with some analytics language added. It says things like “leverages tools built by the Implementation Manager to analyze big data” and asks for proficiency in Python, Spark, or SQL.

That’s pretty far from what I’m actually doing now. We process 1B+ transactions a year, and I’m working on a novel, high-visibility real-time fraud use case for one of our three largest clients. They’re also notoriously difficult to work with.

On the technical side, I’m adding custom platform capabilities for data ingestion and operationalizing ML models, building secure pipelines, doing Spark/PySpark processing, shell automation, SFTP workflows, and building reporting systems that run essentially autonomously. I’m hands-on with almost all of that work, but I’m also project managing the data engineering effort across both companies, coordinating our teams with the client’s technical teams to actually get this stuff into production. Then I’m still doing the analytics on top of the systems I built.

None of my previous responsibilities really went away either. I still manage other technical projects, work directly with the client, and regularly present analytical insights and financial reporting to their executive leadership. I was also heavily involved in work that helped roughly double the size of this client’s contract, which in turn expanded my scope further as we added products and took on more of their transaction volume.

So I’ve ended up doing some weird combination of data science, data/platform engineering, analytics, reporting, project management and client leadership. I actually like the engineering work, so this isn’t a complaint about having to code. I’m more confused about how a role that was originally defined as basically implementation + analytics ended up owning this much production engineering, platform work and client delivery without the classification changing.

The leveling side is where it gets stranger. My boss specifically encouraged me to interview for a Senior Manager opening on our own team. He’d been giving me very positive feedback, telling me he trusted me and that I’d have opportunities, so I went through the full interview process and eventually made it to the VP. During that interview, the VP said something along the lines of, “Well, who else would we hire? You’re already doing the work.” Then later in the conversation he asked whether they’d have to backfill my current position if they promoted me. I ultimately didn’t get the job.

I obviously have no way of knowing whether that question was decisive, but the timing has always bothered me. They left the Senior Manager role open for months and eventually filled the need with another person at the same level I’m currently at rather than hiring a Senior Manager. When I later raised compensation with my boss, the feedback was still positive, but he said I was progressing within the “normal range for my role.”

For additional context, I’m at about $129k base / $148k total comp and I’m already well below the midpoint of the salary band for my existing role. That’s part of what made me start looking harder at whether the role itself is even classified correctly.

For people who have been around senior DS/data organizations, would you still consider this a data science/analytics management role, or at this point is it really some form of data/platform engineering or technical data leadership? I’m also curious how you’d interpret the promotion sequence. Am I reading too much into the backfill question, or does this sound like the classic problem of becoming more valuable in your current seat than the company wants you to be somewhere else?


r/datascience 7d ago

Discussion Anyone who's never worked at FAANG or big tech, do you have regrets about it?

Upvotes

Spent the first half of my 20s doing odd jobs and finishing grad school. Since then I’ve been fortunate to work low stress jobs with good work life balance and moderately high pay.
Now that I’m in my early 30s, living in a high cost of living area with a strong job market, and socializing more, I’ve realized I don’t come close to making what FAANG people make. I can’t help but wonder if that’s something I should be chasing or at least exploring. At the same time, I really value how low stress and stable my job is, even though it pays well below market rate.

For those who chose not to go into big tech/FAANG, do you have any regrets?


r/datascience 7d ago

Career | Europe Thoughts on DS at McKinsey/Bain/BCG

Upvotes

Recently received an offer to go work at McK/BCG/Bain as a forward deployed AI Scientist. Interested to hear thoughts on what people think about DS at these consulting companies.
Do they do fun work?
How cutting edge are they?
WLB?

Currently at a large UK Financial services firm. Is this considered an upgrade?


r/datascience 8d ago

Weekly Entering & Transitioning - Thread 31 Aug, 2026 - 07 Sep, 2026

Upvotes

Welcome to this week's entering & transitioning thread! This thread is for any questions about getting started, studying, or transitioning into the data science field. Topics include:

  • Learning resources (e.g. books, tutorials, videos)
  • Traditional education (e.g. schools, degrees, electives)
  • Alternative education (e.g. online courses, bootcamps)
  • Job search questions (e.g. resumes, applying, career prospects)
  • Elementary questions (e.g. where to start, what next)

While you wait for answers from the community, check out the FAQ and Resources pages on our wiki. You can also search for answers in past weekly threads.


r/datascience 8d ago

Discussion What do I do as DS manager

Upvotes

I'm a Data science manager at a mid-sized company. Every week I have a bunch of meetings, and my reports do all the data science work. They even manage their sprints.

My question is, what do I do? I have free time in my calendar I dont know what to do with. It feels like I'm a viewer in a movie


r/datascience 7d ago

Education Join Kata - a new type to learn SQL

Thumbnail
Upvotes

r/datascience 9d ago

Discussion Do you use a whiteboard when thinking?

Upvotes

Hello all, here is a chill post.

When I was an undergrad, I really liked working things out on a whiteboard. Drawing stuff, talking through ideas out loud, testing little hypotheses.

Now I work in radar DSP, and a lot of my work is code, numerical experiments, deep learning and waiting for training to finish 😅

I’m wondering how other people bring that whiteboard style of thinking into DSP, data science or ML work.

Do you still use a whiteboard regularly, or do you mostly go straight from idea to code?


r/datascience 14d ago

Career | US Feeling frustrated as a junior who has never worked with other analysts or had a senior analyst to learn from.

Upvotes

I've started my career in nonprofits and only worked in nonprofits until now. 3 times now, I have ended up in roles where I am the ONLY analyst on the team. Everyone I work with is either data adjacent, or not an analyst at all. I'm the only person ever working on analytics work, and I have no real life gauge/context on how to do things better in a real world context. I google things all the time, I take courses, but the advice is too generalized and doesn't go deep enough. I need people I can bounce off of. My biggest hope starting as an early career data analyst was that I'd be able to learn from other analyst and fill the gaps in my education with knowledge from mentors.

Instead, I have people looking to me to be an expert in analytics just because I'm the only one available(as if I'm not a junior). Very few opportunities to learn from actual analysts and get experience from them instead of the generalized advice online. I feel like I'm being stunted, but its incredibly hard for me to find roles that are placed in analytics teams, or where I'll be working under a senior analyst (and not just a VP or project manager). Have I screwed myself? Why is it seemingly harder to find analytic roles that work with other analysts?


r/datascience 15d ago

Weekly Entering & Transitioning - Thread 24 Aug, 2026 - 31 Aug, 2026

Upvotes

Welcome to this week's entering & transitioning thread! This thread is for any questions about getting started, studying, or transitioning into the data science field. Topics include:

  • Learning resources (e.g. books, tutorials, videos)
  • Traditional education (e.g. schools, degrees, electives)
  • Alternative education (e.g. online courses, bootcamps)
  • Job search questions (e.g. resumes, applying, career prospects)
  • Elementary questions (e.g. where to start, what next)

While you wait for answers from the community, check out the FAQ and Resources pages on our wiki. You can also search for answers in past weekly threads.


r/datascience 17d ago

AI Man vs Machine (vs Wizard vs Troll) - Article on AI for Game Design

Upvotes

Article

I'm a table top game designer that used AI to build playtesting models. I previously wrote How to Train Your AI Dragon and The Artificial "Intelligence" of Artificial Intelligence which some people here might have read

I really enjoyed writing those articles so decided to enter the writing competition at King's College London to write even more about Machine Learning in game design

My usual style is to mix humor with technical content. I actually wanted to look at the more philosophical side of AI. Looking at why model outputs model, what sort of outcomes AI cannot model, and whether AI actually accomplishes anything important


r/datascience 18d ago

ML New open source relational benchmark and foundation model

Upvotes

New oss relational learning benchmark, leaderboard and TabPFN harness

  1. RelArena-α: standardized relational machine learning model benchmarking
  2. TabPFN-Rel: a harness for tabular foundation model TabPFN-3 for predictions over relational data
  3. RPI-α (Relational Prediction Interface): an interface to run any RelArena model on your own database

- RelArena is open sourced here: https://github.com/PriorLabs/relarena

- Full report: https://arxiv.org/abs/2608.16319

- There's also a higher-level summary of the release by Prior Labs: https://priorlabs.ai/blog-posts/introducing-relarena?utm_source=socials&utm_campaign=relational


r/datascience 18d ago

ML Help point me in the right direction: How to account for decision support systems affecting future training data

Upvotes

I feel like I am googling everything but the exact term I need, and would appreciate someone pointing me in the right direction.

Say you have a customer churn model. You predict a customer has a high likelihood of churning, and then the customer service team gets an alert to intervene. Great! Your model helped mitigate a loss and contributed real value. This is where most tutorials or blog posts on models like this end.

But overtime, customers that have all the signals of churning begin to out perform their expected value.... which would screw up your training data. You've succeeded in putting your thumb on the scale, but in the process potentially damaged the viability of your model.

What is the technical term for this phenomenon? Feedback? It's not target leakage I don't think. Googling "customer churn feedback" just gets you articles about using customer feedback forms as a predictor of churn, which isn't what I want.

Thanks!


r/datascience 19d ago

Tools Interactive tool for learning ML System design for free

Upvotes

I know data scientists are increasingly being asked to own models end-to-end. The problem I see (specially with juniors) is that they jump straight into Docker or other MLOps tools without building the foundation first.

I’ve been in data science for 8+ years and I think the best way to start is with ML system design.

I’ve gone through different resources over the years like Chip Huyen's "Designing Machine Learning Systems" and realized that learning system design just by reading a book or staring at diagrams is really hard. You don't really get the intuition until you can see how the components actually connect and behave.

So I built a free interactive tool based on a real system I deployed. It walks you through how the system was architected so you can build the intuition to design one yourself.

Here it is: https://futureproofds.com/tools/ml-system-map

A few things you can do with it:

  1. Play a flow and watch it run step by step (training, serving a live prediction, the nightly batch, a drift alert firing)

  2. Click any component to see how it works and why it matters in the system

  3. Follow the build order to see how the whole system comes together at different stages

Hoping some of you find it useful. Would love to hear what you think, and let me know if there's any functionality you want me to add.


r/datascience 20d ago

Career | Europe Another rant like interview experience

Upvotes

I was given a home assignment to do modeling for some adtech data. They had no explicit ask about what kind of model or how deep you have to go. Just data and they asked we want to see the modeling.

I spent lot of time in understanding the data, identifying the features, creating labels etc. When it came to modeling I picked Catboost since they handle categorical features quite well. I even mentioned how this can be further tuned and/or different models can be compared. I put it explicitly in a section for future work. Finally this was the thing that got me rejected.

Basically they expected me to compare different model families from more complex deep models to such boosting models. I have worked in this domain and actually such models (catboost) works quite well. You don't need very complex models. I remember in one of the previous jobs, they had like ensemble of 3 deep models which was super slow and was so painful to maintain. I basically replaced that with a boosting model + some probability calibration which did quite well. Also the data size I got for the task isn't big enough to justify such huge models.

In any case, I wish these tasks would be more explicit in what they are looking for. I know they also want to see how I handle ambiguity but it's really hard to assess which side of it is worth handling since I am not building a full fledged system. I explained all the decisions I made and why I did it. Also what I didn't do and why.


r/datascience 21d ago

Discussion How does one prepare for such interviews?

Upvotes

I see posts like these on my Linkedin feed every day. At this juncture, I am not sure if this is true or just one of those AI Slops - I am assuming there's a grain of truth in them.

But now, when I am preparing for interviews and job hunting, I don't think I could have ever imagined answering it in this way, unless I have worked on specific/adjacent use cases.

How does one prepare for such questions?


r/datascience 22d ago

Weekly Entering & Transitioning - Thread 17 Aug, 2026 - 24 Aug, 2026

Upvotes

Welcome to this week's entering & transitioning thread! This thread is for any questions about getting started, studying, or transitioning into the data science field. Topics include:

  • Learning resources (e.g. books, tutorials, videos)
  • Traditional education (e.g. schools, degrees, electives)
  • Alternative education (e.g. online courses, bootcamps)
  • Job search questions (e.g. resumes, applying, career prospects)
  • Elementary questions (e.g. where to start, what next)

While you wait for answers from the community, check out the FAQ and Resources pages on our wiki. You can also search for answers in past weekly threads.


r/datascience 24d ago

ML What Hugging Face learned from reproducing 2,200 ICML papers

Thumbnail
huggingface.co
Upvotes

r/datascience 25d ago

Discussion How widely is R still used in industry today?

Upvotes

I’m a Data Science student (career changer, not in a data related role). My program is focused more on the applied statistics side, so most of my classes use R. I’m already familiar with Python since it was the main language used in my prerequisite courses, and I’ve completed projects using Python, so I’m comfortable with the syntax.

However, I’m really enjoying using and learning R in my classes and seeing what it can do. Many of the statistics textbooks I’m interested in use R as well. I’m starting to explore R more deeply on my own and plan to start using it for personal projects.

But I’m curious, is R still used in industry? I know it’s heavily used in academia. I also know that in the current AI/ML world, Python is used heavily, which is the main reason I use it for all of my personal projects at the moment.

I’d like to eventually be comfortable with both and take advantage of the strengths of each language. But, of course, there are also people who say learning R is a waste of time.