AI도 ‘눈치’ 보나…성능 좋을수록 다수의견에 쏠린다 [후암동 논문 연구소] 작성일 08-18 61 목록 <div id="layerTranslateNotice" style="display:none;"></div> <div class="article_view" data-translation-body="true" data-tiara-layer="article_body" data-tiara-action-name="본문이미지확대_클릭"> <section dmcf-sid="1MTgiLHlH6"> <figure class="figure_frm origin_fig" contents-hash="4fce18580a29269fbe845632de001c6d2af4513540523011d57a28f1395da3ec" dmcf-pid="tRyanoXSH8" dmcf-ptype="figure"> <p class="link_figure"><img alt="사진은 기사 내용과 관련 없음. [게티이미지뱅크]" class="thumb_g_article" data-org-src="https://t1.daumcdn.net/news/202608/18/ned/20260818201145100seql.jpg" data-org-width="1280" dmcf-mid="PTkM5tJ65L" dmcf-mtype="image" height="auto" src="https://img2.daumcdn.net/thumb/R658x0.q70/?fname=https://t1.daumcdn.net/news/202608/18/ned/20260818201145100seql.jpg" width="658"></p> <figcaption class="txt_caption default_figure"> 사진은 기사 내용과 관련 없음. [게티이미지뱅크] </figcaption> </figure> <p contents-hash="a197b524241021f8b84a2f3bfa4f5ef12bbc7cf212b869d11c9a87a4c94ebe1b" dmcf-pid="FeWNLgZvt4" dmcf-ptype="general">[헤럴드경제=장윤우 기자] 인공지능(AI) 에이전트를 한자리에 모아놓자 아무도 시키지 않았는데 스스로 다수 의견을 따라가며 하나로 뭉치는 것으로 나타났다. 성능이 좋은 모델일수록 더 큰 집단에서도 합의에 도달했고 일부는 1000개가 넘는 집단에서 의견을 모았다.</p> <p contents-hash="74e57f5528418b1805c7111e3e06b83eae177698b3d0ef0ce81abfaf7d005401" dmcf-pid="3dYjoa5T5f" dmcf-ptype="general">사람이 서로 아는 사이로 지낼 수 있는 규모가 150~300명이라는 점을 감안하면 사람의 한계를 넘어선 셈이다.</p> <p contents-hash="50c7c5888013309ed34fc504b430f649eda3e73e289b03e60237a18f0c88ab2c" dmcf-pid="0PlneJWIYV" dmcf-ptype="general">최근 국제학술지 사이언스 어드밴시스(Science Advances) 제12권 제33호에 독일 콘스탄츠대 조르다노 데마르초 연구원과 다비트 가르시아 교수 연구팀은 이 같은 내용을 담은 연구 결과를 발표했다. 이탈리아 엔리코페르미연구센터와 오스트리아 복잡계과학허브 연구진이 함께 참여했다.</p> <figure class="figure_frm origin_fig" contents-hash="3d5b78c9044f82c22d1587694e2bc69c7612124edc43f0b6f66082a713269f37" dmcf-pid="pQSLdiYC12" dmcf-ptype="figure"> <p class="link_figure"><img alt="[게티이미지뱅크]" class="thumb_g_article" data-org-src="https://t1.daumcdn.net/news/202608/18/ned/20260818201146695llup.png" data-org-width="860" dmcf-mid="HjuhsIb05M" dmcf-mtype="image" height="auto" src="https://img1.daumcdn.net/thumb/R658x0.q70/?fname=https://t1.daumcdn.net/news/202608/18/ned/20260818201146695llup.png" width="658"></p> <figcaption class="txt_caption default_figure"> [게티이미지뱅크] </figcaption> </figure> <div contents-hash="d09624d62a010f30794b7a46277773a2c34b511a0bc80f6e572e562bce9d6ab9" dmcf-pid="UxvoJnGhX9" dmcf-ptype="general"> 성능 조금만 올라도 몇 배씩 차이 났다 </div> <p contents-hash="5478d45a711ecf816c4214e71480248649b2eba64a2fe4f27dcf1368a31daa75" dmcf-pid="uMTgiLHlGK" dmcf-ptype="general">연구팀은 GPT와 클로드, 라마 계열의 거대언어모델(LLM) 10종으로 각각 AI 에이전트 집단을 만들고 두 가지 의견 중 하나를 무작위로 배정했다. 한 번에 하나씩 골라 나머지 전원이 어떤 의견을 가졌는지 목록으로 보여준 뒤 다시 고르게 하는 방식이었다. 어느 쪽이 옳다는 정보도, 합의하라는 지시도 주지 않았다.</p> <p contents-hash="dd7c8e63004c12e77e3f38e1fab3feafa4961cf7367d744524c1abd45fe262fc" dmcf-pid="7RyanoXSHb" dmcf-ptype="general">의견에는 특정 문장을 사용하지 않고 ‘k’와 ‘z’라는 알파벳을 썼다. ‘예’와 ‘아니오’처럼 뜻이 담긴 단어를 쓰면 LLM이 긍정적인 쪽으로 쏠리는 편향이 강하게 나타나기 때문이다.</p> <p contents-hash="7db522e0a802c22616be87ee3c1ff304ef88963563328200e1a64b54d0d6320f" dmcf-pid="zeWNLgZvXB" dmcf-ptype="general">50개의 에이전트로 실험하자 클로드 3 오푸스와 GPT-4 터보는 20번 모두 전원이 같은 의견으로 모였고, 클로드 3 하이쿠와 GPT-3.5 터보는 한 번도 합의에 이르지 못했다.</p> <figure class="figure_frm origin_fig" contents-hash="8bdeb9247955f621898d590ea9c2882a8e0bb9a446d6637c3e2f6033af6e63a9" dmcf-pid="qdYjoa5T1q" dmcf-ptype="figure"> <p class="link_figure"><img alt="[게티이미지뱅크]" class="thumb_g_article" data-org-src="https://t1.daumcdn.net/news/202608/18/ned/20260818201146929kuma.png" data-org-width="594" dmcf-mid="X53sDr71tx" dmcf-mtype="image" height="auto" src="https://img3.daumcdn.net/thumb/R658x0.q70/?fname=https://t1.daumcdn.net/news/202608/18/ned/20260818201146929kuma.png" width="658"></p> <figcaption class="txt_caption default_figure"> [게티이미지뱅크] </figcaption> </figure> <p contents-hash="d92129c689aa856a97000d50e543b3c973a63fd89e1d77c0efdc3e14f9ffd6a2" dmcf-pid="BJGAgN1yHz" dmcf-ptype="general">연구팀은 AI 모델들이 합의에 이를 수 있는 집단 크기의 한계도 계산했다. 라마 3 70B는 약 30개, GPT-4o는 약 80개, GPT-4 터보는 약 1000개였다. 클로드 3.5 소네트는 1000개 집단에서도 한계가 나타나지 않아 그보다 클 것으로 추정됐다.</p> <p contents-hash="611a6cee0a9bded2fddf61241e16a0e97907a84da87ab7ca634ad2c84e79f2d5" dmcf-pid="biHcajtWX7" dmcf-ptype="general">57개 과목의 지식과 추론 능력을 평가하는 MMLU 점수와 비교하자 상관계수가 0.82로 나왔다. 시험 점수가 높은 모델일수록 더 큰 집단에서도 의견을 하나로 모았다는 뜻이다. 다만 다른 평가 지표에서는 상관관계가 뚜렷하지 않았다.</p> <p contents-hash="1ec27f22a6d0aabc0701e726b8f090940310092f60bbe5d25d6b4ff8298f490f" dmcf-pid="KnXkNAFYZu" dmcf-ptype="general">주목할 점은 성능이 조금 오를 때 한계 크기는 몇 배씩 뛰었다. MMLU 점수 70점대 초반 모델은 집단 크기가 두 자릿수만 돼도 흩어졌지만, 80점대 후반 모델은 1000개를 넘겼다.</p> <p contents-hash="bceb61a2572c78fc2b4e869ec5b5062c1cfba2e53cfa103bf209f5f6d972cee5" dmcf-pid="9LZEjc3GXU" dmcf-ptype="general">실험에 쓰인 모델은 2024년에 공개된 구형 AI로 현재 더 성능이 좋아진 AI의 경우 집단 크기가 어디까지 커질지는 예측하기 어렵다</p> <figure class="figure_frm origin_fig" contents-hash="fa801c246a7feae35f2e02ab21318bbaf7f7a1b97fe7f4ec11651960744d6695" dmcf-pid="2LZEjc3G1p" dmcf-ptype="figure"> <p class="link_figure"><img alt="[AI 생성 이미지]" class="thumb_g_article" data-org-src="https://t1.daumcdn.net/news/202608/18/ned/20260818201147189ecnl.png" data-org-width="1280" dmcf-mid="ZlIdxRvmtQ" dmcf-mtype="image" height="auto" src="https://img3.daumcdn.net/thumb/R658x0.q70/?fname=https://t1.daumcdn.net/news/202608/18/ned/20260818201147189ecnl.png" width="658"></p> <figcaption class="txt_caption default_figure"> [AI 생성 이미지] </figcaption> </figure> <div contents-hash="5853ba68ea62c2a6b1a700babdf0fef2170bef5a0e7d72588772d0cb8125291a" dmcf-pid="Vo5DAk0H50" dmcf-ptype="general"> AI는 관련 없는 과학 이론과 동일한 움직임을 보였다 </div> <p contents-hash="c91081e3329ea845073a0f2845b589b34cc9a740d5ef05529ad3dbfefe2624a0" dmcf-pid="fg1wcEpXY3" dmcf-ptype="general">연구팀은 다수를 따라가는 정도를 ‘다수의 힘’이라는 수치로 나타냈다. 이 값이 크면 다수 쪽으로 쏠리고 작으면 각자 동전을 던지듯 무작위로 고른다.</p> <p contents-hash="d93c0ce5adbdb66c16f608a7e434a9e84a372c13c94d21cf7149a809b5720513" dmcf-pid="4atrkDUZGF" dmcf-ptype="general">에이전트 150개로 시작한 집단은 합의에 실패할 때마다 의견에 따라 쪼개졌고 30개 안팎으로 작아지자 더 이상 갈라지지 않았다. 연구팀은 이 과정이 자석 속 원자들이 같은 방향으로 정렬되는 현상을 설명하는 물리학 모형과 똑같은 수식으로 기술된다고 밝혔다.</p> <figure class="figure_frm origin_fig" contents-hash="15ad5484701ee92722ca7f3605c6b169100df5a76aff40da10d891dc332424c3" dmcf-pid="8NFmEwu55t" dmcf-ptype="figure"> <p class="link_figure"><img alt="[AFP]" class="thumb_g_article" data-org-src="https://t1.daumcdn.net/news/202608/18/ned/20260818201147459cwtq.png" data-org-width="649" dmcf-mid="51du3poMYP" dmcf-mtype="image" height="auto" src="https://img1.daumcdn.net/thumb/R658x0.q70/?fname=https://t1.daumcdn.net/news/202608/18/ned/20260818201147459cwtq.png" width="658"></p> <figcaption class="txt_caption default_figure"> [AFP] </figcaption> </figure> <p contents-hash="403f31f663ce37db5b693ca03242be41fd1846868d9c39e6f2db77c56f8a4a62" dmcf-pid="6j3sDr71t1" dmcf-ptype="general">연구팀은 에이전트들이 그저 다수가 택했다는 이유로 한쪽으로 쏠렸다는 점은 문제가 될 수 있다고 설명했다.</p> <p contents-hash="12e5a0d7399bdfa587b8cc5fc2a25bcc3c7a417645f99cb45fc65a5fb4091613" dmcf-pid="PA0OwmztX5" dmcf-ptype="general">여러 AI가 함께 코드를 짜는 상황에서 이미 널리 쓰인다는 이유만으로 비효율적인 코드나 설계가 그대로 굳어질 수 있다는 것이다.</p> <p contents-hash="7c7618924bd4aae6fa8e1a91227a6a0cc0a1cc747bd0c75053e758757a25a57c" dmcf-pid="QcpIrsqFGZ" dmcf-ptype="general">다만 의견을 뭉치는 능력 자체가 나쁜 것은 아니고 중앙에서 지시하지 않아도 수천 개의 에이전트가 알아서 협업하는 소프트웨어 개발 같은 작업이 가능할 수 있다고 덧붙였다.</p> <p contents-hash="fcec283b0a56f53fc3c86083656c20c8e81cd1ec31287966c4bcecd594310daa" dmcf-pid="xkUCmOB3ZX" dmcf-ptype="general">연구팀은 AI가 왜 다수를 따르는지에 관한 이유로 학습 데이터에 집단행동을 다룬 글이 섞여 있었을 가능성, 사람의 피드백으로 훈련받으며 협조적으로 굴도록 학습됐을 가능성 등을 거론했지만 확인되지 않았다고 한계를 밝혔다.</p> <p contents-hash="d04b86495b059d05cc49c9ccba5a9fc35cc5ff01c9d12b139abbd505d05ed9f5" dmcf-pid="y7AfK2wa1H" dmcf-ptype="general">또한 이번 실험이 극도로 단순화된 상황이라는 점도 인정했다. 둘 중 하나를 고르는 문제였고 정답이 없었으며 AI에 보상도 없었다. 연구팀은 다수를 따르는 것은 가장 단순한 조율 방식일 뿐 상대의 의도를 읽거나 전략적으로 판단하는 진짜 사회적 지능과는 다르다고 밝혔다.</p> <div contents-hash="cd0b757d9d1d9ec016e600a417f91c4a0aae32e435390ed6f8ddc482d870bc80" dmcf-pid="Wzc49VrNtG" dmcf-ptype="general"> 참고논문 </div> <p contents-hash="feda0b1ce144443c562ed036db19df0664c7b54aed800c7d2a89fb59f2026eb6" dmcf-pid="YatrkDUZZY" dmcf-ptype="general">DOI : 10.1126/sciadv.aea6091</p> <p contents-hash="24751a8cbefea56d3b721e2e38b5b0d065b1b28a7067c9dfc26d7d73445e4f2b" dmcf-pid="GNFmEwu5XW" dmcf-ptype="general">논문 정보 : Giordano De Marzo et al. ,AI agents can coordinate via majority-following beyond human scale.Sci. Adv.12,eaea6091(2026).</p> </section> </div> <p class="" data-translation="true">Copyright © 헤럴드경제. 무단전재 및 재배포 금지.</p> 관련자료 이전 '상금에 창업 기회까지' AI 혁신 챌린지…충청권 예선은 천안서 9월 18일 08-18 다음 "'숏폼 영상 자동제작' 뷰컷, 와디즈 펀딩해요" 08-18 댓글 0 등록된 댓글이 없습니다. 로그인한 회원만 댓글 등록이 가능합니다.