πŸ”₯ Matt Dancho (Business Science) πŸ”₯ Profile picture
May 10, 2023 β€’ 8 tweets β€’ 7 min read β€’ Read on X
Learning data science on your own is tough...

...(ahem, it took me 6 years)

So here's some help.

5 Free Books to Cut Your Time In HALF.

Let's go! 🧡

#datascience #rstats #R Image
1. Mastering #Spark with #R

This book solves an important problem- what happens when your data gets too big?

For example, analyzing 100,000,000 time series.

You can do it in R with the tools covered in this book.

Website: therinspark.com Image
2. Geocomputation with #R

Interested in #Geospatial Analysis?

This book is my go-to resource for all things geospatial.

This book covers:
-Making Maps
-Working with Spatial Data
-Applications (Transportation, Geomarketing)

Website: r.geocompx.org Image
3. Tidy Finance with #R

What tools exist in R for #Finance?
And how do I use them?

Answers to these questions are covered in this book!

P.S.- This book uses my R package, #tidyquant

Website: tidy-finance.org Image
4. Text Mining with R

This is a fantastic introduction to text analysis and text mining with the #tidytext R package.

This book singlehandedly made me MORE CONFIDENT with text analysis.

Website: tidytextmining.com Image
5. #Forecasting Principles and Practice

This is the best β€œtheory” book on #timeseries analysis and forecasting.

Topics Covered:
- ARIMA,
- Exponential Smoothing,
- TimeSeries Decomposition
- A lot more!

Website: otexts.com/fpp3/ Image
1-Dollar Bonus Book:

This is a massive value- Gives you a complete plan for EVERYTHING you need to know about learning data science.

It's only a buck.

And it will cut 2-3 years off your journey.

Website: learn.business-science.io/if-i-had-to-le… Image
Want even more help becoming a 6-figure data scientist?

I have a free workshop that will help you become a $100K+ earner as a #DataScientist even in a Recession.

πŸ‘‰Register Here: us02web.zoom.us/webinar/regist… Image

β€’ β€’ β€’

Missing some Tweet in this thread? You can try to force a refresh
γ€€

Keep Current with πŸ”₯ Matt Dancho (Business Science) πŸ”₯

πŸ”₯ Matt Dancho (Business Science) πŸ”₯ Profile picture

Stay in touch and get notified when new unrolls are available from this author!

Read all threads

This Thread may be Removed Anytime!

PDF

Twitter may remove this content at anytime! Save it as PDF for later use!

Try unrolling a thread yourself!

how to unroll video
  1. Follow @ThreadReaderApp to mention us!

  2. From a Twitter thread mention us with a keyword "unroll"
@threadreaderapp unroll

Practice here first or read more on our help page!

More from @mdancho84

Dec 27
🚨 BREAKING: IBM launches a free Python library that converts ANY document to data

Introducing Docling. Here's what you need to know: 🧡 Image
1. What is Docling?

Docling is a Python library that simplifies document processing, parsing diverse formats β€” including advanced PDF understanding β€” and providing seamless integrations with the gen AI ecosystem. Image
2. Document Conversion Architecture

For each document format, the document converter knows which format-specific backend to employ for parsing the document and which pipeline to use for orchestrating the execution, along with any relevant options. Image
Read 8 tweets
Dec 20
Google just dropped a masterclass on Agents.

Here's what's covered in the 54 page PDF: Image
Here's what they cover:

1. From models to agents

2. What an AI Agent is

3. Agentic problem-solving loop (5 steps)

4. Taxonomy of agentic systems (levels 0–4)

5. Core architecture decisions

6. Multi-agent patterns (design patterns)
7. Deployment & Agent Ops (GenAIOps)

8. Interoperability: humans, agents, and payments

9. Security, identity, and governance at scale

10. Learning, self-evolution, and Agent Gym
Read 7 tweets
Dec 17
This 277-page PDF unlocks the secrets of Large Language Models.

Here's what's inside: 🧡 Image
Chapter 1 introduces the basics of pre-training.

This is the foundation of large language models, and common pre-training methods and model architectures will be discussed here. Image
Chapter 2 introduces generative models, which are the large language models we commonly refer to today.

After presenting the basic process of building these models, you explore how to scale up model training and handle long texts. Image
Read 10 tweets
Dec 16
Stanford just made fine-tuning irrelevant with a single paper.

It’s called Agentic Context Engineering (ACE) and it proves you can make models smarter without touching a single weight.

Key takeaways (and get the 23 page PDF): Image
Stanford just released a 23 page paper on Agentic Context Enginnering to improve Agents. Key ideas:

1. ACE = Agentic Context Engineering: treat system prompts + agent memory as a living playbook. Image
2. Log agent trajectories; reflect to extract actionable bullets: strategies, tool schemas, failure modes.

3. Merge updates as append-only deltas with periodic semantic de-duplication to keep the playbook clean.
Read 9 tweets
Dec 8
🚨BREAKING: New Python library for agentic data processing and ETL with AI

Introducing DocETL.

Here's what you need to know: Image
1. What is DocETL?

It's a tool for creating and executing data processing pipelines, especially suited for complex document processing tasks.

It offers:

- An interactive UI playground
- A Python package for running production pipelines Image
2. DocWrangler

DocWrangler helps you iteratively develop your pipeline:

- Experiment with different prompts and see results in real-time
- Build your pipeline step by step
- Export your finalized pipeline configuration for production use Image
Read 8 tweets
Dec 8
🚨 BREAKING: Microsoft launches a free Python library that converts ANY document to Markdown

Introducing Markitdown. Let me explain. 🧡 Image
1. Document Parsing Pipelines

MarkItDown is a lightweight Python utility for converting various files to Markdown for use with LLMs and related text analysis pipelines. Image
2. Supported Documents

MarkItDown supports:

- PDF
- PowerPoint
- Word
- Excel
- Images (EXIF metadata and OCR)
- Audio (EXIF metadata and speech transcription)
- HTML
- Text-based formats (CSV, JSON, XML)
- ZIP files (iterates over contents)
- Youtube URLs
- EPubs Image
Read 10 tweets

Did Thread Reader help you today?

Support us! We are indie developers!


This site is made by just two indie developers on a laptop doing marketing, support and development! Read more about the story.

Become a Premium Member ($3/month or $30/year) and get exclusive features!

Become Premium

Don't want to be a Premium member but still want to support us?

Make a small donation by buying us coffee ($5) or help with server cost ($10)

Donate via Paypal

Or Donate anonymously using crypto!

Ethereum

0xfe58350B80634f60Fa6Dc149a72b4DFbc17D341E copy

Bitcoin

3ATGMxNzCUFzxpMCHL5sWSt4DVtS8UqXpi copy

Thank you for your support!

Follow Us!

:(