View a markdown version of this page

SageMaker Training Jobs에서 Nova 2.0의 지도 미세 조정(SFT) - Amazon Nova

SageMaker Training Jobs에서 Nova 2.0의 지도 미세 조정(SFT)

사전 조건

훈련 작업을 시작하기 전에 다음을 확인하세요.

  • 입력 데이터와 훈련 작업 출력을 저장하는 Amazon S3 버킷. 두 가지 데이터를 하나의 버킷에 함께 저장하거나, 각 데이터 유형을 별도의 버킷에 저장할 수 있습니다. 버킷이 모든 훈련 관련 리소스를 생성하는 AWS 리전과 동일한 리전에 있는지 확인합니다. 자세한 내용은 범용 버킷 생성을 참조하세요.

  • 훈련 작업 실행 권한이 있는 IAM 역할. IAM 정책을 AmazonSageMakerFullAccess에 연결해야 합니다. 자세한 내용은 How to use SageMaker AI execution roles를 참조하세요.

  • 기본 Amazon Nova 레시피, Amazon Nova 레시피 가져오기 참조.

SFT란 무엇인가요?

지도 미세 조정(SFT)은 레이블이 지정된 입력 및 출력 페어를 사용하여 언어 모델을 훈련합니다. 모델은 프롬프트와 응답으로 구성된 데모 예제로부터 학습하여 특정 태스크, 지침 또는 원하는 동작에 맞게 해당 기능을 개선합니다.

SFT가 사용 사례에 적합한지 확인하려면 지도 미세 조정(SFT) 섹션을 참조하세요.

훈련 작업 시작

데이터 준비

데이터 형식, 지원되는 기능, 제약 조건 및 SFT 훈련 데이터 준비 모범 사례에 대한 자세한 내용은 SFT on Amazon Nova 2용 데이터 준비하기 섹션을 참조하세요.

데이터 업로드

데이터세트는 SageMaker Training Jobs에서 액세스할 수 있는 버킷에 업로드해야 합니다. 올바른 권한 설정에 대한 자세한 내용은 사전 조건을 참조하세요.

하이퍼파라미터 선택 및 레시피 업데이트

Nova 2.0의 설정은 Nova 1.0의 설정과 거의 동일합니다. 입력 데이터가 S3에 업로드되면 미세 조정 폴더 GitHub의 SageMaker HyperPod 레시피에서 설명하는 레시피를 사용합니다. Nova 2.0 Lite의 경우 다음은 사용 사례에 따라 업데이트할 수 있는 몇 가지 주요 하이퍼파라미터입니다. 다음은 Nova 2.0 Lite SFT PEFT 레시피에 대한 예제입니다. 컨테이너 이미지 URI의 경우 708977205387.dkr.ecr.us-east-1.amazonaws.com/nova-fine-tune-repo:SM-TJ-SFT-V2-latest를 사용하여 SFT 미세 조정 작업을 실행합니다.

샘플 입력

run: name: {peft_recipe_job_name} model_type: amazon.nova-2-lite-v1:0:256k model_name_or_path: {peft_model_name_or_path} data_s3_path: {train_dataset_s3_path} # SageMaker HyperPod (SMHP) only and not compatible with SageMaker Training jobs. Note replace my-bucket-name with your real bucket name for SMHP job replicas: 4 # Number of compute instances for training, allowed values are 4, 8, 16, 32 output_s3_path: "" # Output artifact path (Hyperpod job-specific; not compatible with standard SageMaker Training jobs). Note replace my-bucket-name with your real bucket name for SMHP job training_config: max_steps: 10 # Maximum training steps. Minimal is 4. save_steps: 10 # How many training steps the checkpoint will be saved. Should be less than or equal to max_steps save_top_k: 1 # Keep top K best checkpoints. Note supported only for SageMaker HyperPod jobs. Minimal is 1. max_length: 32768 # Sequence length (options: 8192, 16384, 32768 [default], 65536) global_batch_size: 32 # Global batch size (options: 32, 64, 128) reasoning_enabled: true # If data has reasoningContent, set to true; otherwise False lr_scheduler: warmup_steps: 15 # Learning rate warmup steps. Recommend 15% of max_steps min_lr: 1e-6 # Minimum learning rate, must be between 0.0 and 1.0 optim_config: # Optimizer settings lr: 1e-5 # Learning rate, must be between 0.0 and 1.0 weight_decay: 0.0 # L2 regularization strength, must be between 0.0 and 1.0 adam_beta1: 0.9 # Exponential decay rate for first-moment estimates, must be between 0.0 and 1.0 adam_beta2: 0.95 # Exponential decay rate for second-moment estimates, must be between 0.0 and 1.0 peft: # Parameter-efficient fine-tuning (LoRA) peft_scheme: "lora" # Enable LoRA for PEFT lora_tuning: alpha: 64 # Scaling factor for LoRA weights ( options: 32, 64, 96, 128, 160, 192), lora_plus_lr_ratio: 64.0

이 레시피에는 Nova 1.0과 거의 동일한 하이퍼파라미터도 포함되어 있습니다. 주목할 만한 하이퍼파라미터는 다음과 같습니다.

  • max_steps - 작업을 실행하려는 단계 수입니다. 일반적으로 하나의 에포크(전체 데이터세트를 통해 하나가 실행됨)의 경우 단계 수 = 데이터 샘플 수/글로벌 배치 크기입니다. 단계 수가 많고 글로벌 배치 크기가 작을수록 작업 실행 시간이 길어집니다.

  • reasoning_enabled - 데이터세트의 추론 모드를 제어합니다. 옵션:

    • true: 추론 모드 활성화(높은 노력의 추론과 동일)

    • false: 추론 모드 비활성화

    참고: SFT의 경우 추론 노력 수준을 세밀하게 제어하지 않습니다. reasoning_enabled: true를 설정하면 전체 추론 기능이 활성화됩니다.

  • peft.peft_scheme - 이를 'lora'로 설정하면 PEFT 기반 미세 조정이 활성화됩니다. null(따옴표 없음)로 설정하면 전체 순위 미세 조정이 활성화됩니다.

훈련 작업 시작

from sagemaker.pytorch import PyTorch # define OutputDataConfig path if default_prefix: output_path = f"s3://{bucket_name}/{default_prefix}/{sm_training_job_name}" else: output_path = f"s3://{bucket_name}/{sm_training_job_name}" output_kms_key = "<KMS key arn to encrypt trained model in Amazon-owned S3 bucket>" # optional, leave blank for Amazon managed encryption recipe_overrides = { "run": { "replicas": instance_count, # Required "output_s3_path": output_path }, } estimator = PyTorch( output_path=output_path, base_job_name=sm_training_job_name, role=role, disable_profiler=True, debugger_hook_config=False, instance_count=instance_count, instance_type=instance_type, training_recipe=training_recipe, recipe_overrides=recipe_overrides, max_run=432000, sagemaker_session=sagemaker_session, image_uri=image_uri, output_kms_key=output_kms_key, tags=[ {'Key': 'model_name_or_path', 'Value': model_name_or_path}, ] ) print(f"\nsm_training_job_name:\n{sm_training_job_name}\n") print(f"output_path:\n{output_path}")
from sagemaker.inputs import TrainingInput train_input = TrainingInput( s3_data=train_dataset_s3_path, distribution="FullyReplicated", s3_data_type="Converse", ) estimator.fit(inputs={"validation": val_input}, wait=False)
참고

Nova 2.0 Lite의 지도 미세 조정에 대해서는 검증 데이터세트 전달이 지원되지 않습니다.

작업을 시작하는 방법:

  • 데이터세트 경로 및 하이퍼파라미터로 레시피 업데이트

  • 노트북에서 지정된 셀을 실행하여 훈련 작업 제출

노트북은 작업 제출을 처리하고 상태 추적을 제공합니다.