ScalaSparkStreamingJobProps

class aws_cdk.aws_glue.ScalaSparkStreamingJobProps(*, role, script, connections=None, continuous_logging=None, default_arguments=None, description=None, glue_version=None, job_name=None, max_concurrent_runs=None, max_retries=None, security_configuration=None, tags=None, timeout=None, enable_metrics=None, enable_observability_metrics=None, spark_ui=None, worker_configuration=None, class_name, extra_files=None, extra_jars=None, extra_jars_first=None, job_run_queuing_enabled=None)

Bases: SparkJobProps

Properties for creating a Scala Spark ETL job.

Parameters:
  • role (IRole) – IAM Role (required) IAM Role to use for Glue job execution Must be specified by the developer because the L2 doesn’t have visibility into the actions the script(s) takes during the job execution The role must trust the Glue service principal (glue.amazonaws.com) and be granted sufficient permissions.

  • script (Code) – Script Code Location (required) Script to run when the Glue job executes. Can be uploaded from the local directory structure using fromAsset or referenced via S3 location using fromBucket

  • connections (Optional[Sequence[IConnection]]) – Connections (optional) List of connections to use for this Glue job Connections are used to connect to other AWS Service or resources within a VPC. Default: [] - no connections are added to the job

  • continuous_logging (Union[ContinuousLoggingProps, Dict[str, Any], None]) – Enables continuous logging with the specified props. Default: - continuous logging is enabled.

  • default_arguments (Optional[Mapping[str, str]]) – Default Arguments (optional) The default arguments for every run of this Glue job, specified as name-value pairs. This map is the escape hatch for Glue job arguments that this construct does not model. It MUST NOT be used to set arguments that already have a dedicated prop — configure those through the corresponding prop instead (continuousLogging, enableMetrics, enableObservabilityMetrics, sparkUI, className, extraJars, extraJarsFirst, extraPythonFiles, extraFiles). Passing a construct-managed argument (e.g. --enable-continuous-cloudwatch-log, --enable-metrics, --enable-spark-ui, --job-language) or a Glue-reserved argument (--debug, --mode, --JOB_NAME, --endpoint) here throws at synthesis time, so there is exactly one way to express each intent. Also note that these are emitted verbatim into the CloudFormation template, so avoid placing secrets here in plaintext. Pass secrets to the job at runtime through AWS Secrets Manager instead. A synthesis-time warning is emitted when an argument key looks like a credential and holds a plaintext literal. Default: - no arguments

  • description (Optional[str]) – Description (optional) Developer-specified description of the Glue job. Default: - no value

  • glue_version (Optional[GlueVersion]) – Glue Version The version of Glue to use to execute this job. Default: - determined by the job type: 4.0 for ETL and Streaming, 5.0 for Flex, 3.0 for Python Shell

  • job_name (Optional[str]) – Name of the Glue job (optional) Developer-specified name of the Glue job. Default: - a name is automatically generated

  • max_concurrent_runs (Union[int, float, None]) – Max Concurrent Runs (optional) The maximum number of runs this Glue job can concurrently run. An error is returned when this threshold is reached. The maximum value you can specify is controlled by a service limit. Default: 1

  • max_retries (Union[int, float, None]) – Max Retries (optional) Maximum number of retry attempts Glue performs if the job fails. Default: 0

  • security_configuration (Optional[ISecurityConfiguration]) – Security Configuration (optional) Defines the encryption options for the Glue job. Default: - no security configuration.

  • tags (Optional[Mapping[str, str]]) – Tags (optional) A list of key:value pairs of tags to apply to this Glue job resources. Default: {} - no tags

  • timeout (Optional[Duration]) – Timeout (optional) The maximum time that a job run can consume resources before it is terminated and enters TIMEOUT status. Specified in minutes. Default: - no value set; Glue applies its service default (2880 minutes for non-streaming jobs)

  • enable_metrics (Optional[bool]) – Enable profiling metrics for the Glue job. When enabled, adds ‘–enable-metrics’ to job arguments. Default: true

  • enable_observability_metrics (Optional[bool]) – Enable observability metrics for the Glue job. When enabled, adds ‘–enable-observability-metrics’: ‘true’ to job arguments. Default: true

  • spark_ui (Union[SparkUIProps, Dict[str, Any], None]) – Enables the Spark UI debugging and monitoring with the specified props. Default: - Spark UI debugging and monitoring is disabled.

  • worker_configuration (Union[WorkerConfiguration, Dict[str, Any], None]) – The worker type and the number of workers allocated when a job runs. Default: - the job runs with the G_1X worker type and 10 workers.

  • class_name (str) – Class name (required for Scala scripts) Package and class name for the entry point of Glue job execution for Java scripts.

  • extra_files (Optional[Sequence[Code]]) – Additional files, such as configuration files that AWS Glue copies to the working directory of your script before executing it. Default: - no extra files specified.

  • extra_jars (Optional[Sequence[Code]]) – Extra Jars S3 URL (optional) S3 URL where additional jar dependencies are located. Default: - no extra jar files

  • extra_jars_first (Optional[bool]) – Setting this value to true prioritizes the customer’s extra JAR files in the classpath. Default: false - priority is not given to user-provided jars

  • job_run_queuing_enabled (Optional[bool]) – Specifies whether job run queuing is enabled for the job runs for this job. A value of true means job run queuing is enabled for the job runs. If false or not populated, the job runs will not be considered for queueing. If this field does not match the value set in the job run, then the value from the job run field will be used. This property must be set to false for flex jobs. If this property is enabled, maxRetries must be set to zero. Default: - no job run queuing

ExampleMetadata:

fixture=_generated

Example:

# The code below shows an example of how to instantiate this type.
# The values are placeholders you should change.
import aws_cdk as cdk
from aws_cdk import aws_glue as glue
from aws_cdk import aws_iam as iam
from aws_cdk import aws_logs as logs
from aws_cdk import aws_s3 as s3

# bucket: s3.Bucket
# code: glue.Code
# connection: glue.Connection
# log_group: logs.LogGroup
# role: iam.Role
# security_configuration: glue.SecurityConfiguration

scala_spark_streaming_job_props = glue.ScalaSparkStreamingJobProps(
    class_name="className",
    role=role,
    script=code,

    # the properties below are optional
    connections=[connection],
    continuous_logging=glue.ContinuousLoggingProps(
        enabled=False,

        # the properties below are optional
        conversion_pattern="conversionPattern",
        log_group=log_group,
        log_stream_prefix="logStreamPrefix",
        quiet=False
    ),
    default_arguments={
        "default_arguments_key": "defaultArguments"
    },
    description="description",
    enable_metrics=False,
    enable_observability_metrics=False,
    extra_files=[code],
    extra_jars=[code],
    extra_jars_first=False,
    glue_version=glue.GlueVersion.V0_9,
    job_name="jobName",
    job_run_queuing_enabled=False,
    max_concurrent_runs=123,
    max_retries=123,
    security_configuration=security_configuration,
    spark_ui=glue.SparkUIProps(
        bucket=bucket,
        prefix="prefix"
    ),
    tags={
        "tags_key": "tags"
    },
    timeout=cdk.Duration.minutes(30),
    worker_configuration=glue.WorkerConfiguration(
        number_of_workers=123,
        worker_type=glue.WorkerType.STANDARD
    )
)

Attributes

class_name

Class name (required for Scala scripts) Package and class name for the entry point of Glue job execution for Java scripts.

connections

Connections (optional) List of connections to use for this Glue job Connections are used to connect to other AWS Service or resources within a VPC.

Default:

[] - no connections are added to the job

continuous_logging

Enables continuous logging with the specified props.

Default:
  • continuous logging is enabled.

See:

https://docs.aws.amazon.com/glue/latest/dg/aws-glue-programming-etl-glue-arguments.html

default_arguments

Default Arguments (optional) The default arguments for every run of this Glue job, specified as name-value pairs.

This map is the escape hatch for Glue job arguments that this construct does not model. It MUST NOT be used to set arguments that already have a dedicated prop — configure those through the corresponding prop instead (continuousLogging, enableMetrics, enableObservabilityMetrics, sparkUI, className, extraJars, extraJarsFirst, extraPythonFiles, extraFiles). Passing a construct-managed argument (e.g. --enable-continuous-cloudwatch-log, --enable-metrics, --enable-spark-ui, --job-language) or a Glue-reserved argument (--debug, --mode, --JOB_NAME, --endpoint) here throws at synthesis time, so there is exactly one way to express each intent.

Also note that these are emitted verbatim into the CloudFormation template, so avoid placing secrets here in plaintext. Pass secrets to the job at runtime through AWS Secrets Manager instead. A synthesis-time warning is emitted when an argument key looks like a credential and holds a plaintext literal.

Default:
  • no arguments

See:

https://docs.aws.amazon.com/glue/latest/dg/aws-glue-programming-etl-glue-arguments.html for a list of reserved parameters

description

Description (optional) Developer-specified description of the Glue job.

Default:
  • no value

enable_metrics

Enable profiling metrics for the Glue job.

When enabled, adds ‘–enable-metrics’ to job arguments.

Default:

true

enable_observability_metrics

Enable observability metrics for the Glue job.

When enabled, adds ‘–enable-observability-metrics’: ‘true’ to job arguments.

Default:

true

extra_files

Additional files, such as configuration files that AWS Glue copies to the working directory of your script before executing it.

Default:
  • no extra files specified.

See:

https://docs.aws.amazon.com/glue/latest/dg/aws-glue-programming-etl-glue-arguments.html

extra_jars

Extra Jars S3 URL (optional) S3 URL where additional jar dependencies are located.

Default:
  • no extra jar files

extra_jars_first

Setting this value to true prioritizes the customer’s extra JAR files in the classpath.

Default:

false - priority is not given to user-provided jars

See:

--user-jars-first in https://docs.aws.amazon.com/glue/latest/dg/aws-glue-programming-etl-glue-arguments.html

glue_version

Glue Version The version of Glue to use to execute this job.

Default:
  • determined by the job type: 4.0 for ETL and Streaming, 5.0 for Flex, 3.0 for Python Shell

job_name

Name of the Glue job (optional) Developer-specified name of the Glue job.

Default:
  • a name is automatically generated

job_run_queuing_enabled

Specifies whether job run queuing is enabled for the job runs for this job.

A value of true means job run queuing is enabled for the job runs. If false or not populated, the job runs will not be considered for queueing. If this field does not match the value set in the job run, then the value from the job run field will be used. This property must be set to false for flex jobs. If this property is enabled, maxRetries must be set to zero.

Default:
  • no job run queuing

max_concurrent_runs

Max Concurrent Runs (optional) The maximum number of runs this Glue job can concurrently run.

An error is returned when this threshold is reached. The maximum value you can specify is controlled by a service limit.

Default:

1

max_retries

Max Retries (optional) Maximum number of retry attempts Glue performs if the job fails.

Default:

0

role

IAM Role (required) IAM Role to use for Glue job execution Must be specified by the developer because the L2 doesn’t have visibility into the actions the script(s) takes during the job execution The role must trust the Glue service principal (glue.amazonaws.com) and be granted sufficient permissions.

See:

https://docs.aws.amazon.com/glue/latest/dg/getting-started-access.html

script

Script Code Location (required) Script to run when the Glue job executes.

Can be uploaded from the local directory structure using fromAsset or referenced via S3 location using fromBucket

security_configuration

Security Configuration (optional) Defines the encryption options for the Glue job.

Default:
  • no security configuration.

spark_ui

Enables the Spark UI debugging and monitoring with the specified props.

Default:
  • Spark UI debugging and monitoring is disabled.

See:

https://docs.aws.amazon.com/glue/latest/dg/aws-glue-programming-etl-glue-arguments.html

tags

value pairs of tags to apply to this Glue job resources.

Default:

{} - no tags

Type:

Tags (optional) A list of key

timeout

Timeout (optional) The maximum time that a job run can consume resources before it is terminated and enters TIMEOUT status.

Specified in minutes.

Default:
  • no value set; Glue applies its service default (2880 minutes for non-streaming jobs)

worker_configuration

The worker type and the number of workers allocated when a job runs.

Default:
  • the job runs with the G_1X worker type and 10 workers.