Skip to content

Augmentation

OCR 학습 시 사용 가능한 augmentation 종류는 다음과 같습니다 (총 10종류)

[
    "vertical_flip",
    "horizontal_flip",
    "rotate",
    "perspective_transform",
    "adjust_brightness",
    "adjust_contrast",
    "adjust_hue",
    "adjust_saturation",
    "adjust_gamma",
    "blur",
]

이를 위한 image augmentation API는 Preview API와 Trainer Config의 2가지로 구성됩니다.

  • Preview API: 유저가 세팅한 파라미터에 대해 이미지 미리보기 제공
  • Trainer Config: 유저가 세팅한 파라미터를 학습에 사용하기 위해 Trainer.build()에 넣어주는 config

ocr.ImageProcessor

Bases: _APIDecorator

Image Augmentation Preview를 위한 API 입니다. 각 augmentation들이 특정 파라미터 값에 대해 이미지를 어떻게 변형 시키는지 확인할 수 있습니다.

Note

일반적으로 학습 시에는 파라미터를 특정 이 아닌 범위로 설정하여 해당 범위에서 매번 랜덤한 값을 선택해 이미지에 적용합니다. 따라서 preview API의 입력 파라미터와 학습 시 넘겨주는 파라미터는 대부분 vs 범위의 차이를 가지게 됩니다. 예를 들어 preview API에서 rotate의 경우 angle (float) 값을 받지만, 학습 config에서는 angle_limit (List[float]) 범위를 받게됩니다. 각 augmentation을 학습에 사용시 필요한 config는 각 함수 설명의 Trainer Config 섹션을 참고하세요.

Usage

rotate augmentation preview 예제입니다. 상세 설명은 각 함수 설명 참고.

image = np.zeros((100, 100, 3), dtype=np.unit8)
error, augmented_image = ImageProcessor.rotate(image=image, angle=15)

vertical_flip(image)

image를 상하로 뒤집습니다.

Parameters:

Name Type Description Default
image np.ndarray

augmentation을 적용할 image 입니다.

data type: uint8
channel: H x W / H x W x 1  - Gray
         H x W x 3 - RGB
         H x W x 4 - RGBA

required

Returns:

Type Description
np.ndarray

np.ndarray: augmentation이 적용된 image 입니다.

Example
error, augmented_image = ImageProcessor.vertical_flip(image=image)
Trainer Config
{
    "_target_": "vertical_flip",
}

horizontal_flip(image)

image를 좌우로 뒤집습니다.

Parameters:

Name Type Description Default
image np.ndarray

augmentation을 적용할 image 입니다.

data type: uint8
channel: H x W / H x W x 1  - Gray
         H x W x 3 - RGB
         H x W x 4 - RGBA

required

Returns:

Type Description
np.ndarray

np.ndarray: augmentation이 적용된 image 입니다.

Example
error, augmented_image = ImageProcessor.horizontal_flip(image=image)
Trainer Config
{
    "_target_": "horizontal_flip",
}

rotate(image, angle=0.0)

image를 angle만큼 회전시킵니다.

Parameters:

Name Type Description Default
image np.ndarray

augmentation을 적용할 image 입니다.

data type: uint8
channel: H x W / H x W x 1  - Gray
         H x W x 3 - RGB
         H x W x 4 - RGBA

required
angle float

회전하는 각도 입니다. 유효 범위는 다음과 같습니다. [-360.0, 360.0]. Defaults to 0.0.

0.0

Returns:

Type Description
np.ndarray

np.ndarray: augmentation이 적용된 image 입니다.

Example
error, augmented_image = ImageProcessor.rotate(image=image, angle=60.0)
Trainer Config
{
    "_target_": "rotate",
    "angle_limit": List[float],  # angle 최소 최대 범위
}

perspective_transform(image, offset_top_left, offset_top_right, offset_bottom_right, offset_bottom_left)

이미지를 투영 변환(Perspective Transform)합니다. image를 다른 각도(시점)에서 바라본 형태로 변환합니다.

Parameters:

Name Type Description Default
image np.ndarray

augmentation을 적용할 image 입니다.

data type: uint8
channel: H x W / H x W x 1  - Gray
         H x W x 3 - RGB
         H x W x 4 - RGBA

required
offset_top_left OffsetType

Perspective Transform 을 위한 도착점의 좌표를 계산할 때 사용됩니다. 원본 이미지의 좌측 상단 모서리를 얼마만큼 중심부로 이동시킬 지에 대한 실수 값이 들어있습니다. 즉, Perspective Transform 의 좌측 상단 도착점 point_dst 는 아래와 같이 계산할 수 있습니다.

x` = 0 + width * offset_top_left[0]
y` = 0 + height * offset_top_left[1]
point_dst = (x`, y`)

offset_* 의 범위는 [0, 49] 로, 각 도착지 점들은 이미지의 중간 선을 지나칠 수 없습니다. 이를 통해 이미지가 반전되는 정도의 왜곡을 방지합니다.

required
offset_bottom_right OffsetType

offset_top_left 와 동일하되 우측 하단 점을 나타냅니다. Perspective Transform 의 우측 하단 도착점 point_dst 는 아래와 같이 계산할 수 있습니다.

x` = width - width * offset_bottom_right[0]
y` = height - height * offset_bottom_right[1]
point_dst = (x`, y`)
required
offset_top_right OffsetType

offset_top_left 와 동일하되 우측 상단 점을 나타냅니다.

required
offset_bottom_left OffsetType

offset_top_left 와 동일하되 좌측 하단 점을 나타냅니다.

required

Returns:

Type Description
np.ndarray

np.ndarray: augmentation이 적용된 image 입니다.

Example
error, augmented_image = ImageProcessor.perspective_transform(
    image=image,
    offset_top_left=(0.25, 0.30),
    offset_top_right=(0.10, 0.40),
    offset_bottom_right=(0.00, 0.20),
    offset_bottom_left=(0.20, 0.35),
)
Trainer Config
{
    "_target_": "perspective_transform",
    "intensity_limit": List[float],  # intensity 최소 최대 범위
}

adjust_brightness(image, brightness=0.0)

image의 밝기 (brightness)를 변경합니다.

Parameters:

Name Type Description Default
image np.ndarray

augmentation을 적용할 image 입니다.

data type: uint8
channel: H x W / H x W x 1  - Gray
         H x W x 3 - RGB
         H x W x 4 - RGBA

required
brightness float

밝기를 변화 강도입니다. 유효 범위는 다음과 같습니다. [-1.00, 1.00]. Defaults to 0.0.

0.0

Returns:

Type Description
np.ndarray

np.ndarray: augmentation이 적용된 image 입니다.

Example
error, augmented_image = ImageProcessor.adjust_brightness(image=image, brightness=0.3)
Trainer Config
{
    "_target_": "adjust_brightness",
    "brightness_limit": List[float],  # brightness 최소 최대 범위
}

adjust_contrast(image, contrast=0.0)

image의 대비 (contrast)를 변경합니다.

Parameters:

Name Type Description Default
image np.ndarray

augmentation을 적용할 image 입니다.

data type: uint8
channel: H x W / H x W x 1  - Gray
         H x W x 3 - RGB
         H x W x 4 - RGBA

required
contrast float

대비 변화 강도입니다. 유효 범위는 다음과 같습니다. [-1.00, 1.00]. Defaults to 0.0.

0.0

Returns:

Type Description
np.ndarray

np.ndarray: augmentation이 적용된 image 입니다.

Example
error, augmented_image = ImageProcessor.adjust_contrast(image=image, contrast=0.7)
Trainer Config
{
    "_target_": "adjust_contrast",
    "contrast_limit": List[float],  # contrast 최소 최대 범위
}

adjust_hue(image, hue=0.0)

image의 색조 (hue)를 변경합니다.

Parameters:

Name Type Description Default
image np.ndarray

augmentation을 적용할 image 입니다.

data type: uint8
channel: H x W / H x W x 1  - Gray
         H x W x 3 - RGB
         H x W x 4 - RGBA

required
hue float

색조 변화 강도입니다. 유효 범위는 다음과 같습니다. [-1.00, 1.00]. Defaults to 0.0.

0.0

Returns:

Type Description
np.ndarray

np.ndarray: augmentation이 적용된 image 입니다.

Example
error, augmented_image = ImageProcessor.adjust_hue(image=image, hue=0.7)
Trainer Config
{
    "_target_": "adjust_hue",
    "hue_limit": List[float],  # hue 최소 최대 범위
}

adjust_saturation(image, saturation=0.0)

image의 채도 (saturation)를 변경합니다.

Parameters:

Name Type Description Default
image np.ndarray

augmentation을 적용할 image 입니다.

data type: uint8
channel: H x W / H x W x 1  - Gray
         H x W x 3 - RGB
         H x W x 4 - RGBA

required
saturation float

채도 변화 강도입니다. 유효 범위는 다음과 같습니다. [-1.00, 1.00]. Defaults to 0.0.

0.0

Returns:

Type Description
np.ndarray

np.ndarray: augmentation이 적용된 image 입니다.

Example
error, augmented_image = ImageProcessor.adjust_saturation(image=image, saturation=0.7)
Trainer Config
{
    "_target_": "adjust_saturation",
    "saturation_limit": List[float],  # saturation 최소 최대 범위
}

adjust_gamma(image, gamma=0.0)

image의 gamma를 조절하여 밝기를 변화시킵니다.

Parameters:

Name Type Description Default
image np.ndarray

augmentation을 적용할 image 입니다.

data type: uint8
channel: H x W / H x W x 1  - Gray
         H x W x 3 - RGB
         H x W x 4 - RGBA

required
gamma float

조절할 gamma value 입니다. 유효 범위는 다음과 같습니다. [-1.0, 1.0]. Defaults to 0.0.

0.0

Returns:

Type Description
np.ndarray

np.ndarray: augmentation이 적용된 image 입니다.

Example
error, augmented_image = ImageProcessor.adjust_gamma(image=image, gamma=0.5)
Trainer Config
{
    "_target_": "adjust_gamma",
    "gamma_limit": List[float],  # gamma 최소 최대 범위
}

blur(image, ksize=1)

image에 averaging blur를 적용합니다.

Parameters:

Name Type Description Default
image np.ndarray

augmentation을 적용할 image 입니다.

data type: uint8
channel: H x W / H x W x 1  - Gray
         H x W x 3 - RGB
         H x W x 4 - RGBA

required
ksize int

blur kernel의 크기 입니다. 유효 범위는 다음과 같습니다. [1, 100]. Defaults to 1.

1

Returns:

Type Description
np.ndarray

np.ndarray: augmentation이 적용된 image 입니다.

Example
error, augmented_image = ImageProcessor.blur(image=image, ksize=3)
Trainer Config
{
    "_target_": "blur",
    "ksize_limit": List[int],  # ksize 최소 최대 범위
}

Trainer Config

Trainer에 넘겨주는 augmentation config는 List[Dict] 형태이며, 적용할 augmentation들의 파라미터(Dict)를 리스트로 모은 것입니다. 해당 리스트를 Trainer.build() config의 augmentation 섹션에 넣어주면 학습 시 augmentation이 적용됩니다.

  • Augmentation config 예시 (10가지 augmentation 모두 사용하는 경우, 파라미터는 유저가 선택한 값)

    augmentation_config = [
        {
            "_target_": "vertical_flip",
        },
        {
            "_target_": "horizontal_flip",
        },
        {
            "_target_": "rotate",
            "angle_limit": [-30, 30],
        },
        {
            "_target_": "perspective_transform",
            "intensity_limit": [0, 30],
        },
        {
            "_target_": "adjust_brightness",
            "brightness_limit": [-0.2, 0.2],
        },
        {
            "_target_": "adjust_contrast",
            "contrast_limit": [-0.2, 0.2],
        },
        {
            "_target_": "adjust_hue",
            "hue_limit": [-0.2, 0.2],
        },
        {
            "_target_": "adjust_saturation",
            "saturation_limit": [-0.2, 0.2],
        },
        {
            "_target_": "adjust_gamma",
            "gamma_limit": [-0.2, 0.2],
        },
        {
            "_target_": "blur",
            "ksize_limit": [3, 7],
        },
    ]
    

  • Augmentation config 예시2 ("vertical_flip"만 사용하는 경우)

    augmentation_config = [
        {
            "_target_": "vertical_flip",
        },
    ]
    

  • Trainer build config 예시

    augmentation_config = [
        {
            "_target_": "vertical_flip",
        },
        {
            "_target_": "rotate",
            "angle_limit": [-30, 30],
        },
    ]
    
    trainer_config = {
        "model": "shallow",
        ...
        "augmentation": augmentation_config,  # augmentation section에 삽입
        ...
    }
    
    error, trainer = Trainer.build(config=trainer_config, data)