AI NEWS

AI 뉴스

국내외 AI 기술 동향과 산업 뉴스를
전문가 시각으로 큐레이션합니다.

AI 뉴스

문장 변환기를 이용한 다중 벡터 임베딩 모델의 학습 및 미세 조정

2026.08.26읽기
AI 뉴스

How loveholidays is making everyone a builder with Codex

Discover how loveholidays uses OpenAI Codex to make software development accessible across the business, helping teams turn ideas into products faster.

2026.08.26읽기
AI 뉴스

Hugging Face 보안 사건과 앞으로 나아갈 길

OpenAI shares findings from the Hugging Face security incident and the steps we’re taking to strengthen AI model security, monitoring, and alignment.

2026.08.26읽기
AI 뉴스

5 ways to upgrade your home decor with Google Search

Learn how to use Google Search tools to find home decor inspiration, shop for furniture, and tackle DIY projects.

2026.08.25읽기
AI 뉴스

Granite 4.2 LLMs: How They're Built

2026.08.25읽기
AI 뉴스

Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original

2026.08.25읽기
AI 뉴스

양자화 인식 복구: 압축된 4비트 모델이 원본의 전체 정밀도 모델보다 우수한 성능을 보여준다

대규모 언어 모델을 축소하는 것은 거의 항상 비용을 수반합니다. 효율적인 배포를 위한 표준 방식은 먼저 아키텍처를 압축하는 것입니다. 레이어, 헤드 또는 뉴런을 제거하여 파라미터 수를 줄인 다음, 남은 가중치를 4비트로 양자화하여 메모리와 연산량을 더욱 줄입니다. 이 두 단계 모두 상당한 절약을 가져오지만, 추론, 수학적 문제 해결 및 코드 생성과 같이 사람들이 실제로 중요하게 생각하는 기능은 체계적으로 저하됩니다. 이러한 이유로, 진지한 배포 파이프라인은 모델을 실제 환경에 배포하기 전에 일반적으로 '힐링'이라고 하는 복구 단계를 추가합니다. gpt-oss , NVIDIA의 Nemotron 제품군, 그리고 저희의 Hypernova 60B 와 같은 최근의 오픈 웨이트 모델들은 모두 이러한 압축 후 복구 방식의 변형을 사용합니다. 저희 최신 논문인 " 양자화 인식 복구: 압축된 4비트 LLM 복구를 위한 실용적인 방법"에서는 기존 학계에서 대부분 미해결 상태로 남아 있던 질문을 던집니다. 즉, 모델이 양자화뿐 아니라 구조적 압축을 이미 거친 경우, 복구 단계는 실제로 얼마나 잘 작동하며, 올바른 복구 방법은 무엇인가 하는 것입니다. 저희는 양자화 인식 복구(QAH)라는 기법을 소개하고, 120비트 GPT-OSS 모델을 60비트 파라미터로 압축하고 MXFP4로 양자화한 후 적용했습니다. 그 결과, 9개의 벤치마크 중 7개에서 자체 원본(bfloat16) 버전보다 우수한 성능을 보이는 모델을 생성했습니다. 4비트 모델은 크기가 더 작고, 실행 비용이 더 저렴하며, 양자화의 원본 체크포인트보다 더 높은 정확도를 보였습니다. 이는 4비트 모델과 그 원본인 16비트 모델 간의 일반적인 관계를 뒤집는 결과입니다.

2026.08.25읽기
AI 뉴스

The full stack behind abundant intelligence

OpenAI CFO Sarah Friar explains how advances across chips, compute, models, and products compound to deliver more useful intelligence at greater scale and lower cost.

2026.08.25읽기
AI 뉴스

Jalapeño’s first results show industry-leading speed and efficiency in AI inference

Jalapeño is a custom inference chip from OpenAI that delivers faster, more power-efficient AI inference, with higher throughput and lower latency for modern models.

2026.08.25읽기
AI 뉴스

Jalapeño의 첫 번째 결과는 AI 추론에서 업계 최고 수준의 속도와 효율성을 보여줍니다.

OpenAI의 첫 번째 맞춤형 추론 칩인 Jalapeño를 발표한 이후, 저희는 해당 칩과 이를 기반으로 구축된 시스템을 테스트해 왔습니다. 그 결과, 상당한 성능 향상을 확인할 수 있었습니다. Jalapeño는 단위 전력당 더 많은 AI 작업을 처리하면서도 응답 속도를 더욱 빠르게 합니다. 기존 하드웨어 시스템에서는 처리량과 지연 시간 중 하나를 선택해야 했던 것과 달리, Jalapeño는 단일 아키텍처로 더 높은 처리량과 더 낮은 지연 시간을 모두 제공합니다.

2026.08.25읽기
AI 뉴스

Introducing the Admin plugin for ChatGPT Work and Codex

Use the Admin plugin for ChatGPT Work and Codex to analyze workspace usage, manage members and permissions, adjust limits, and act on admin requests.

2026.08.25읽기
AI 뉴스

Wire It, Run It, Deploy It: AI Workflows in Gradio

2026.08.25읽기
AI 뉴스

ChatGPT Work 및 Codex용 관리자 플러그인을 소개합니다.

워크스페이스 활동을 파악하고, 액세스 및 사용량을 관리하고, 지원되는 관리자 작업을 하나의 대화에서 수행할 수 있습니다.

2026.08.25읽기
AI 뉴스

Disrupting a new covert influence campaign from Russia

OpenAI banned Russia-origin accounts using AI to promote a fake Israel-based think tank and a “sovereignty” index praising Russia and criticizing the West.

2026.08.25읽기
AI 뉴스

Advancing price-performance for developers with GPT‑5.6 in Kiro

GPT‑5.6 is now available in Kiro, helping developers plan, build, review, and test software with better price-performance.

2026.08.24읽기
AI 뉴스

Measuring benchmark optimization in speech recognition

2026.08.21읽기
AI 뉴스

How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code

2026.08.21읽기
AI 뉴스

Up to 3.2x Faster Inference with LFM2.5-DSpark

2026.08.20읽기
AI 뉴스

Introducing Intelligence Age

Introducing Intelligence Age, a new OpenAI blog exploring how transformative AI could reshape power, governance, the economy, and individual freedom.

2026.08.20읽기
AI 뉴스

Stampli cuts launch hours by 68% using ChatGPT Work

With a fixed deadline and design resources committed elsewhere, Stampli used Codex and ChatGPT Work to compress weeks of launch production into days.

2026.08.20읽기