Apache Spark - A unified analytics engine for large-scale data processing
Star History
Momentum
+1.1K
STARS · LAST 30 DAYS
37
PER DAY
#1
MOST-STARRED Scala
| Window | 7 days | 30 days | 90 days |
|---|---|---|---|
| Stars gained | +259 | +1.1K | +3.3K |
| Per day | 37 | 37 | 37 |
| Forks gained | +147 | +588 | +1.5K |
spark gained 1.1K stars in the last 30 days, about 37 a day, and now has 44.1K. It is about 13 years old and has averaged roughly 3.4K stars a year. It ranks #1 among Scala repositories and #233 across all languages on GitHubRepo.
Trending Record
6
DAYS ON TRENDING
#200
BEST RANK
Sep 23, 2026
FIRST APPEARANCE
Active
STATUS TODAY
spark has maintained a continuous presence across global trending indexes, peaking at #200. Below is the 30-day activity profile:
💡 Overview
spark is an open-source project written in Scala: Apache Spark - A unified analytics engine for large-scale data processing.
Engineered for speed, consistency, and developer ease, it solves common hurdles in big-data, java, jdbc. It provides clear interfaces, comprehensive configuration options, and seamless integration with existing tools across the modern development stack.
⚡ Key Features
Optimized execution pipeline written in Scala for predictable speed.
Zero-friction configuration with comprehensive sensible defaults out of the box.
Cross-platform runtime support across Linux, macOS, and Windows environments.
Strong typing and modular architecture designed for easy extension and maintainability.
Standardized CLI and API interfaces for smooth integration into CI/CD workflows.
Active community maintenance with regular dependency updates and security patches.
📥 Installation
$ git clone https://github.com/apache/spark.git
cd spark
⚙ System Requirements
Platforms
- • macOS
- • Linux
- • Windows
Runtime & Dependencies
Scala environment and standard tooling
Architecture
x86_64, ARM64 (Apple Silicon & Graviton)
🧠 How It Works
spark coordinates its core functionality through a modular Scala pipeline. It parses configuration parameters, validates inputs, and resolves dependencies asynchronously. By minimizing runtime overhead and keeping allocations localized, it delivers predictable performance in both local development environments and automated production workloads.
🎯 Production Use Cases
Production System Integration
Embed spark into Scala backend services to handle core application logic.
CI/CD Automated Pipelines
Run automated validation, builds, and integration suites during deployments.
Developer Tooling & Workflows
Accelerate developer onboarding with pre-configured project utilities.
Open Source Extension
Fork and customize internal modules under the repository's open Apache-2.0 license.
🚀 Getting Started
Install spark using your package manager: `git clone https://github.com/apache/spark.git`
Initialize your project workspace or configuration file for spark.
Import spark into your codebase or invoke it directly from your terminal.
Execute your test suite or run `spark --help` to verify successful setup.
👍 Strengths
⚠️ Considerations
⇄ Alternatives & Direct Competitors
👥 Who Should Use This
Developers and engineering teams building with Scala, seeking reliable, tested, and actively maintained tooling for production workloads.
🏆 Nearby in the Rankings
apache/spark is currently ranked #233 by stars across every repository tracked on GitHubRepo. These are adjacent projects:
| Rank | Repository | Language | Stars | Action |
|---|---|---|---|---|
| #228 | spf13/cobra | Go | ★ 44.7K | Compare ↗ |
| #229 | zen-browser/desktop | C++ | ★ 44.6K | Compare ↗ |
| #230 | sharkdp/fd | Rust | ★ 44.6K | Compare ↗ |
| #231 | juanfont/headscale | Go | ★ 44.2K | Compare ↗ |
| #232 | danielmiessler/Fabric | Go | ★ 44.1K | Compare ↗ |
| #233 | apache/spark This Project | Scala | ★ 44.1K | |
| #234 | colinhacks/zod | TypeScript | ★ 44K | Compare ↗ |
| #235 | ray-project/ray | Python | ★ 43.9K | Compare ↗ |
| #236 | astaxie/build-web-application-with-golang | Go | ★ 43.9K | Compare ↗ |
| #237 | HeyPuter/puter | TypeScript | ★ 43.6K | Compare ↗ |
| #238 | reactive-resume/reactive-resume | TypeScript | ★ 43.4K | Compare ↗ |
Frequently Asked Questions
What does spark do? +
Apache Spark - A unified analytics engine for large-scale data processing
What language is spark written in? +
The primary language is Scala. Topics include: big-data, java, jdbc, python, r.
Is spark actively maintained? +
Yes, the last recorded push was on Sep 26, 2026 with 574 open issues being tracked.
How many stars does spark have? +
spark has 44,052 stars and 29,398 forks on GitHub.
How does spark rank among GitHub repositories? +
With 44,052 stars, apache/spark is ranked #233 globally across all repositories tracked on GitHubRepo and #1 among Scala projects.
What license is spark distributed under? +
The repository reports a Apache-2.0 license. Always verify the repository LICENSE file for legal terms.