Automic VaultAutomic Vault

brew

apache-spark mit Homebrew installieren

Prüfe Installationswege, Executables, Metadaten und Sicherheitshinweise für apache-spark in AI-Agent-Workflows.

Installation

Weitere Installationsbefehle

macOS

Homebrewverifiziert · 100%
brew install apache-spark

local Homebrew formula metadata

Überblick

Paketzusammenfassung

Engine for large-scale data processing

Befehle und Aliase

  • docker-image-tool.sh
  • find-spark-home
  • load-spark-env.sh
  • pyspark
  • run-example
  • spark-beeline
  • spark-class
  • spark-connect-shell
  • spark-pipelines
  • spark-shell
  • spark-sql
  • spark-submit
  • sparkR

Verlauf

Projektgeschichte und Nutzung

Apache Spark is a general-purpose engine for large-scale data processing. For package-manager users, it is the canonical install that gives you `spark-submit`, language shells, SQL tooling, example runners, and runtime scripts for local and cluster-oriented workflows.

Projektgeschichte

Spark originated at the UC Berkeley AMPLab as a faster, more interactive alternative to earlier MapReduce-centered data processing systems. Its project history is closely tied to resilient distributed datasets, in-memory computation, and developer-friendly APIs for Scala, Python, Java, SQL, and R.

Spark became an Apache project and grew into a broad analytics engine rather than a single-purpose batch runner. The official project history notes its Apache Software Foundation path and the release line that made Spark a standard part of the big-data toolchain.

Over time Spark absorbed major adjacent workloads: Spark SQL and DataFrames for structured data, MLlib for machine learning, GraphX for graph processing, Structured Streaming for stream processing, and Spark Connect for client-server connectivity.

Adoptionsgeschichte

Spark's adoption history is unusually deep for a package-manager formula because it crossed from research project to de facto data-platform component. It is used for ETL, interactive analytics, machine learning pipelines, and streaming workloads across local machines, YARN, Mesos-era clusters, Kubernetes, and managed cloud services.

The supplied Homebrew package data shows a CLI-heavy install surface: `spark-submit`, `spark-shell`, `pyspark`, `spark-sql`, `sparkR`, `spark-class`, and helper scripts. That executable set mirrors the way Spark became both an application runtime and a command-line toolbox.

Wie es verwendet wird

The main package workflow is submitting applications with `spark-submit`, opening interactive shells with `spark-shell` or `pyspark`, running SQL through `spark-sql`, and configuring behavior through files in `$SPARK_HOME/conf`.

Spark users often install it locally even when production jobs run elsewhere, because the local CLI is useful for testing jobs, validating dependencies, developing notebooks or scripts, and matching cluster runtime behavior.

Warum Paket-Nerds sich dafür interessieren

Spark is a classic heavyweight formula: it is mostly scripts plus a large JVM distribution, but those scripts define the ergonomics of a whole data ecosystem. Packagers care about Java compatibility, Python/R bindings, shell wrappers, classpaths, examples, and config file layout.

It is also one of the packages that turns a laptop into a miniature data platform. A formula install can run local mode, submit to clusters, or serve as a client for remote compute, which makes it more than a simple CLI utility.

Zeitleiste

  • 2009: Spark begins at UC Berkeley AMPLab.
  • 2010: Spark is open sourced.
  • 2013: Spark enters the Apache Incubator.
  • 2014: Spark becomes an Apache top-level project.
  • 2020s: Spark continues expanding SQL, streaming, Kubernetes, and client-server features.

Related projects

  • Apache Hadoop and YARN are central to Spark's early cluster deployment history.
  • Apache Hive influenced Spark SQL's data-warehouse compatibility story.
  • Delta Lake, Apache Iceberg, and Apache Hudi are common table-format companions in modern Spark deployments.

Sicherheitslage

Risikostufe: yellow

broad file, network, media, or database tool signal. generalized runtime or code generation signal.

Risikoklassifikator

yellow Risiko · mittel Konfidenz · runtime

Warum

  • broad file, network, media, or database tool signal
  • generalized runtime or code generation signal

Signale

  • text:shell
  • text:sql,image

Installationsverhalten

  • In den Formelmetadaten ist kein Homebrew-Post-install-Hook erfasst.
  • Homebrew-Bottle-Metadaten sind für 1 Plattformziele verfügbar.
  • Installiert mit 1 Laufzeitabhängigkeiten.

Empfohlene Prüfung

Prüfe vor unbeaufsichtigter Agent-Nutzung, ob das Tool Klartext-Credentials liest, Remote-Zustand schreibt, Artefakte veröffentlicht oder Plugins ausführt.

local files

Configuration and credential file locations

These source-backed paths show where this package keeps local settings or durable credentials. Automic Vault can use them as review targets for secret scanning, migration, and command approval.

Configuration files

Config paths the tool may read or write during local use.

Unix
$SPARK_HOME/conf/spark-defaults.conf$SPARK_HOME/conf/spark-env.sh$SPARK_HOME/conf/log4j2.properties

Executables

Installierte Executables

BefehlArtSichtbarkeitHinweis
docker-image-tool.shcliglobales Executable
find-spark-homecliglobales Executable
load-spark-env.shcliglobales Executable
pysparkcliglobales Executable
run-examplecliglobales Executable
spark-beelinecliglobales Executable
spark-classcliglobales Executable
spark-connect-shellcliglobales Executable
spark-pipelinescliglobales Executable
spark-shellcliglobales Executable
spark-sqlcliglobales Executable
spark-submitcliglobales Executable
sparkRcliglobales Executable

Aktualität

Version und Aktualität

Diese Signale trennen das Alter der Seitengenerierung, Aktivität des Paketmanagers und Upstream-Release-Vergleich. Versionsrückstand wird nur gemeldet, wenn eine Evidenz-URL und vergleichbare Versionen vorhanden sind.

Seite generiert2026-07-25
Manager-Version4.2.0
Manager aktualisiert2026-07-15
lokale DatenOK
Upstreamnot checked
neueste erkannte Versionnicht erkannt

https://spark.apache.org/

Installationsmetadaten

Paketmetadaten

Paketschlüsselbrew:apache-spark
Version4.2.0
PaketmanagerHomebrew
Paketmanager-Seitehttps://formulae.brew.sh/formula/apache-spark
Homepagehttps://spark.apache.org/
Repositoryhttps://github.com/apache/spark
Upstream-Dokumentationhttps://spark.apache.org/docs/latest
LizenzApache-2.0
Quellarchivhttps://www.apache.org/dyn/closer.lua?path=spark/spark-4.2.0/spark-4.2.0-bin-hadoop3.tgz
Zuletzt aktualisiert2026-07-15T03:07:47Z
Pulseupdated
Abhängigkeitenopenjdk@21
Bottleverfügbar (auf all)
Homebrew post-installnicht definiert
Dienstkeiner deklariert

Registry-Fakten

Details aus der Quelldatenbank

Source DatabaseHomebrew formula API
Taphomebrew/core
Full Nameapache-spark
Version Scheme0
Revision0
Head VersionHEAD
Bottle Stable Root URLhttps://ghcr.io/v2/homebrew/core
Deprecatedno
Disabledno
Keg Onlyno
URL Keys
  • head
  • stable

Quellspur

Aus Repository-Daten generiert

Diese Seite wird von av-web aus dem privaten Paket-SQLite-Artefakt bereitgestellt, das scripts/generate-pkg-sqlite.py erstellt.

Verwendete Quellen

  • Geiger risk classifier
  • Nucleus package database
  • av.db category and tag curation
  • cross-ecosystem install command graph
  • curated configuration and credential file locations
  • curated package history
  • package relationship graph
  • package version freshness
  • package-page enrichment