0. 개요
이 포스트에서는 현재 입력되는 DTO 간의 중복 탐지 로직의 코드를 개선하고, 10만 건 입력 시 응답 시간이 오래 걸리는 원인을 발견한 후 이를 해결해 나가고자 한다.
1. 현재 코드 살펴보기
// 아이템 이상 탐지 및 저장 서비스 메서드 public List<Item> createCommonItem(List<CreateCommonItemDocumentReqDto> reqDtos, File savedFile) { // 1. 검증 및 중복 매핑 결과 취득 DuplicateValidationResult validationResult = itemDocumentDuplicateValidator.markDuplicatesForCommon( reqDtos, itemRepository.findAllByDeletedAtIsNullOrderByIdAsc() ); List<Item> existingItemsToUpdate = new ArrayList<>(); List<Boolean> isDuplicateFlags = new ArrayList<>(); // 2. DTO -> Item 엔티티 및 DuplicatedGroup 연관관계 구성 List<Item> itemsToSave = processItemsAndGroups( reqDtos, validationResult, savedFile, existingItemsToUpdate, isDuplicateFlags ); // 3. 기타 이상 탐지 및 이슈(Issue) 수집 List<Issue> issues = detectIssues(itemsToSave, reqDtos, isDuplicateFlags); // 4. 데이터 일괄 저장 return saveAllEntities(itemsToSave, existingItemsToUpdate, issues); }
ItemDocumentDuplicateValidator에서 중복 row 를 탐지하고, 해당 정보를DuplicateValidationResult에 담아서 전달하고 있는 모습이다.
public DuplicateValidationResult markDuplicatesForCommon( List<CreateCommonItemDocumentReqDto> dtos, List<Item> allExistingItems ) { if (dtos.isEmpty()) { return new DuplicateValidationResult(Map.of(), List.of()); } // DB 기존 데이터 Key -> Item 매핑 Map<String, Item> existingDbMap = new HashMap<>(); for (Item item : allExistingItems) { String key = generateKey(...); existingDbMap.putIfAbsent(key, item); } // 현재 입력 데이터 처리 (DB 체크 + 자가 중복 체크가 혼재) Map<String, CreateCommonItemDocumentReqDto> firstSeenMap = new HashMap<>(); for (CreateCommonItemDocumentReqDto dto : dtos) { String normalizedName = itemNameMapper.map(dto.getRawItemName()); dto.setNormalizedItemName(normalizedName); String key = generateKey(...); // 모든 DTO에 대해 Key 생성 if (existingDbMap.containsKey(key) || firstSeenMap.containsKey(key)) { dto.setDuplicateGroupKey(key); if (firstSeenMap.containsKey(key) && firstSeenMap.get(key).getDuplicateGroupKey() == null) { firstSeenMap.get(key).setDuplicateGroupKey(key); } } else { firstSeenMap.put(key, dto); } } return new DuplicateValidationResult(existingDbMap, dtos); }
- 하나의 메서드에서 "DTO 간 자가 중복 체크" 와 "DB 데이터 포함 중복 체크" 를 함께 진행하고 있어 역할이 혼재되어 있다.
- 모든 DTO 에 대해
generateKey()를 실행하고 있기 때문에, 중복이 의심되는 경우에만 Key 를 생성하도록 수정이 필요하다.
2. 구조화 해보기
기존의 markDuplicatesForCommon 메서드에서 DTO 체크와 DB 체크가 얽혀 있던 것을 역할에 따라 메서드로 분리한다.
public DuplicateValidationResult markDuplicatesForCommon( List<CreateCommonItemDocumentReqDto> dtos, List<Item> allExistingItems ) { if (dtos.isEmpty()) { return new DuplicateValidationResult(Map.of(), List.of()); } // 1. DTO 품목명 정규화 (선행 조건) normalizeItemNames(dtos); // 2. 파일 내부 (DTO 간) 중복 탐지 및 GroupKey 부여 markFileSelfDuplicates(dtos); // 3. DB 기존 데이터와의 중복 탐지 Map<String, Item> existingDbMap = markDbDuplicates(dtos, allExistingItems); return new DuplicateValidationResult(existingDbMap, dtos); }
private void normalizeItemNames(List<CreateCommonItemDocumentReqDto> dtos) { for (CreateCommonItemDocumentReqDto dto : dtos) { String normalizedName = itemNameMapper.map(dto.getRawItemName()); dto.setNormalizedItemName(normalizedName); } }
private void markFileSelfDuplicates(List<CreateCommonItemDocumentReqDto> dtos) { // String Key 대신 DTO 자체를 Key로 사용 → generateKey() 호출 최소화 Map<CreateCommonItemDocumentReqDto, CreateCommonItemDocumentReqDto> firstSeenMap = new HashMap<>(135_000); for (CreateCommonItemDocumentReqDto dto : dtos) { CreateCommonItemDocumentReqDto firstSeenDto = firstSeenMap.get(dto); if (firstSeenDto == null) { // 최초 등장 → Map 에 등록 firstSeenMap.put(dto, dto); } else { // 중복 발견 → 필요 시점에만 String Key를 1회 생성하여 공유 if (firstSeenDto.getDuplicateGroupKey() == null) { String groupKey = generateKeyFromDto(firstSeenDto); firstSeenDto.setDuplicateGroupKey(groupKey); } dto.setDuplicateGroupKey(firstSeenDto.getDuplicateGroupKey()); } } }
private Map<String, Item> markDbDuplicates( List<CreateCommonItemDocumentReqDto> dtos, List<Item> allExistingItems ) { if (allExistingItems == null || allExistingItems.isEmpty()) return Map.of(); Map<String, Item> existingDbMap = new HashMap<>(allExistingItems.size()); for (Item item : allExistingItems) { existingDbMap.putIfAbsent(generateKeyFromItem(item), item); } for (CreateCommonItemDocumentReqDto dto : dtos) { // 자가 중복 Key 가 이미 있으면 재활용, 없으면 신규 생성 String dtoKey = (dto.getDuplicateGroupKey() != null) ? dto.getDuplicateGroupKey() : generateKeyFromDto(dto); if (existingDbMap.containsKey(dtoKey)) { dto.setDuplicateGroupKey(dtoKey); } } return existingDbMap; }
- DTO 체크 / DB 체크를 메서드별로 분리하여 SRP (단일 책임 원칙) 준수
- 기존
String Key 기반 Map→ DTO 객체 자체를 Key로 사용하여generateKey()호출을 중복이 감지된 시점에만 실행
3. 소요 시간 측정
시간 소요가 많은 부분을 파악하기 위해 System.nanoTime() 을 이용하여 단계별로 측정하였다.
3-1. 중복 탐지 소요 시간 (10만 건 · DB 데이터 없는 경우)
전체 응답 — 약 6분 소요

단계별 측정 결과

당연히 중복 탐지 때문에 오래 걸렸을 것이라 생각했는데, 파싱 + 정규화 + 중복 탐지는 약 7초에 불과했다! (정규화 프로세스가 가장 많이 소요)
따라서 다른 부분의 소요 시간도 측정해보기로 했다.
3-2. 서비스 단 소요 시간 측정
1,000 건
[서비스] 2. 연관 관계 구성: 487.06 ms
[서비스] 3. 이상 탐지 및 이슈 수집: 112.38 ms
[서비스] 4. DB 커밋: 396.85 ms
[서비스] 5. 총 시간: 1526.00 ms
[파일서비스] 0-1. 파서 추출: 1.00 ms
[파일서비스] 0-2. file DB 입력: 466.84 ms
[파일서비스] 0-3. 파싱 수행: 130.75 ms
[파일서비스] 0-4. itemService 수행: 1530.46 ms
10만 건
[서비스] 2. 연관 관계 구성: 3436.36 ms
[서비스] 3. 이상 탐지 및 이슈 수집: 3226.39 ms
[서비스] 4. DB 커밋: 16982.73 ms
[서비스] 5. 총 시간: 26595.20 ms
[파일서비스] 0-1. 파서 추출: 9.16 ms
[파일서비스] 0-2. file DB 입력: 610.82 ms
[파일서비스] 0-3. 파싱 수행: 285410.31 ms ← 병목!
[파일서비스] 0-4. itemService 수행: 26599.11 ms
파싱 시간이 전체의 대부분을 차지한다는 것을 확인했다. 소요 시간이 큰 순서대로 성능 개선을 진행한다.
4. 파싱 시간 개선하기
csvReader.readAll() 로 인한 메모리 폭증
List<String[]> rows = csvReader.readAll();로 전체 파일을 한번에 메모리에 올려 사용량이 급증한다.- Full GC 가 수초간 발생하여 CPU 점유도 높아진다.
- 메모리 생성 최소화가 필요하다.
List<String[]> rows = csvReader.readAll(); // ← 전체를 한번에 메모리에 적재 if (rows.isEmpty()) return list; for (int i = 1; i < rows.size(); i++) { String[] row = rows.get(i); // ... DTO 생성 list.add(dto); }
List<CreateCommonItemDocumentReqDto> list = new ArrayList<>(135_000); String[] row; long rowNo = 0; csvReader.readNext(); // 헤더 스킵 // readAll() 대신 한 줄씩 읽기 while ((row = csvReader.readNext()) != null) { rowNo++; ParseCsvValueHelper.ParseContext context = new ParseCsvValueHelper.ParseContext(); CreateCommonItemDocumentReqDto dto = CreateCommonItemDocumentReqDto.builder() .rowNo(rowNo) .docId(ParseCsvValueHelper.parseString(getValue(row, 0))) .sourceType(ParseCsvValueHelper.parseString(getValue(row, 1))) .supplierName(ParseCsvValueHelper.parseString(getValue(row, 2))) .rawItemName(ParseCsvValueHelper.parseString(getValue(row, 3))) .spec(ParseCsvValueHelper.parseString(getValue(row, 4))) .unit(ParseCsvValueHelper.parseString(getValue(row, 5))) .priceBefore(ParseCsvValueHelper.parseLong(getValue(row, 6), context)) .priceAfter(ParseCsvValueHelper.parseLong(getValue(row, 7), context)) .effectiveDate(ParseCsvValueHelper.parseDate(getValue(row, 8), context)) .hasParseError(context.hasError()) .build(); list.add(dto); }
개선 성능 측정 결과
[서비스] 5. 총 시간: 33170.79 ms
[파일서비스] 0-4. itemService 수행: 33177.03 ms
개선 후 응답 결과

추가 개선 — Batch Size 별 파싱 + 저장 (Consumer 방식)
리스트에 모든 데이터를 담지 않고 정해진 Batch Size 별로 파싱 후 DB에 저장하는 방식이다.
public void parseAndConsume(MultipartFile file, Consumer<List<CreateCommonItemDocumentReqDto>> batchConsumer) { int batchSize = 1000; List<CreateCommonItemDocumentReqDto> buffer = new ArrayList<>(batchSize); try (CSVReader csvReader = ...) { String[] row; long rowNo = 0; csvReader.readNext(); // 헤더 스킵 while ((row = csvReader.readNext()) != null) { rowNo++; // ... DTO 생성 후 buffer 에 추가 buffer.add(dto); // 1,000개 모이면 즉시 저장 후 버퍼 비우기 if (buffer.size() >= batchSize) { batchConsumer.accept(buffer); buffer.clear(); } } if (!buffer.isEmpty()) batchConsumer.accept(buffer); } } // ── 서비스 레이어에서 사용 ── csvParser.parseAndConsume(file, batchList -> { documentRepository.batchInsert(batchList); // JdbcTemplate.batchUpdate 실행 });
List<> 에 모든 입력 데이터를 담고 중복 로직을 수행하는 방식을 함께 수정해야 적용이 가능하다.
List 기반 중복 로직을 개선해 나가면서 적용해볼 예정이다.
5. 앞으로 개선 사항
[서비스] 2. 연관 관계 구성: 3947.84 ms
[서비스] 3. 이상 탐지 및 이슈 수집: 6501.81 ms
[서비스] 4. DB 커밋: 18339.77 ms
[서비스] 5. 총 시간: 33170.79 ms
참고
하지만 아래 문제가 발생할 수 있다고 판단하여 적용하지 않았다.
- 변하지 않아야 하는 DB 데이터의 무결성이 깨진다.
- 여러 클라이언트가 동시에 파일을 올릴 경우 데이터들 간의 정합성이 깨질 수 있다.
'프로젝트 > 보살핌 프로젝트' 카테고리의 다른 글
| 2. 입력 최적화 (0) | 2026.08.31 |
|---|---|
| 1. 이상 탐지 리팩토링 (0) | 2026.08.27 |
| 0. 해커톤 개요 (0) | 2026.08.25 |






































