MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
We introduce MLE-bench, a benchmark for measuring how well AI agents perform at machine learning engineering.
AI NEWS
국내외 AI 기술 동향과 산업 뉴스를
전문가 시각으로 큐레이션합니다.
We introduce MLE-bench, a benchmark for measuring how well AI agents perform at machine learning engineering.
Hearst’s iconic brands bring curated lifestyle and local news content to OpenAI’s products.
In addition to securing $6.6 billion in new funding from leading investors, we have established a new $4 billion credit facility with leading banks, including JPMorgan Chase, Citi, Goldman Sachs, Morgan Stanley, Santander, Wells Fargo, SMBC, UBS, and HSBC.
We are making progress on our mission to ensure that artificial general intelligence benefits all of humanity.
Developers can now build fast speech-to-speech experiences into their applications
Developers can now fine-tune GPT-4o with images and text to improve vision capabilities
Offering automatic discounts on inputs that the model has recently seen
Fine-tune a cost-efficient model with the outputs of a large frontier model–all on the OpenAI platform
Altera uses GPT-4o to build a new area of human collaboration
OpenAI banned accounts that accessed its models through an Israel-based startup to generate conversations and send links to gambling sites on X.