/instascrape

Flexible, lightweight Python 3 Instagram scraper designed for data mining

Primary LanguagePythonMIT LicenseMIT

instascrape logo

Instagram scraping for humans

What is it?

instascrape is a powerful, lightweight library for scraping Instagram data with no configurations necessary! It is designed with flexibility and developer productivity in mind so you can stop wasting valuable time collecting data and just start analyzing πŸ’ͺ

Version Language Code style: black Release License

Downloads Activity Dependencies Issues Size

Example showing tech profile scrapes

Key features

  • πŸ’ͺ Powerful, object-oriented scraping tools as well as a variety of useful functions
  • πŸ’ƒ Flexibly determines whether you want to scrape HTML, JSON, BeautifulSoup, or request and scrape the URL itself
  • πŸ’Ύ Download content to your computer as png, jpg, mp4, and mp3
  • 🎼 Expressive and consistent API for concise and elegant code
  • πŸ“Š Designed for seamless integration with Selenium, Pandas, and other industry standard tools for data collection and analysis
  • πŸ”¨ Lightweight: you don't have to build a hammer factory when all you need is the hammer
  • πŸ•ΈοΈ The only hard dependencies are Requests and Beautiful Soup; no more worrying about configurations or webdrivers
  • ⌚ Proven to work as of October 29, 2020

Table of Contents


πŸ’» Installation

Minimum Python version

This library currently requires Python 3.7 or higher.

pip

Install from PyPI using

$ pip3 install insta-scrape

WARNING: make sure you install insta-scrape and not a package with a similar name!


πŸ”Ž Sample Usage

All top-level, ready-to-use features can be imported using:

from instascrape import *

instascrape uses clean, consistent, and expressive syntax to make the developer experience as painless as possible.

# Instantiate the scraper objects 
google = Profile('https://www.instagram.com/google/')
google_post = Post('https://www.instagram.com/p/CG0UU3ylXnv/')
google_hashtag = Hashtag('https://www.instagram.com/explore/tags/google/')

# Load their respective data 
google.load()
google_post.load()
google_hashtag.load()

After being scraped, relevant attributes can be accessed with dot (.) or bracket ([]) notation

print(google.followers)
print(google_post['hashtags'])
print(google_hashtag.amount_of_posts)
>>> 12262794
>>> ['growwithgoogle']
>>> 9053408

πŸ“š Documentation

The official documentation can be found on Read The Docs πŸ“°


πŸ“° Blog Posts

Check out blog posts on DEV for ideas and tutorials!


πŸ™ Contributing

All contributions, bug reports, bug fixes, documentation improvements, enhancements, and ideas are welcome!

Feel free to open an Issue or look at existing Issues to get a dialogue going on what you want to see added/changed/fixed.

Beginners to open source are highly encouraged to participate and ask questions ❀️


πŸ•ΈοΈ Dependencies

Instascrape primarily relies on two third-party libraries for requesting and scraping Instagram HTML content:

  1. Requests: HTTP requests
  2. BeautifulSoup: Scraping and parsing HTML data.

The rest of its functionality is provided directly from Python 3's standard library for unobtrusive code under the hood with little to no overhead.


πŸ’³ License

MIT


❔ Support

Reach out to me if you have questions or ideas!