xAI introduces Grok, its AI chatbot connected to X
xAI has introduced Grok, an AI chatbot with access to real-time information through X and a personality shaped by humor. The tool arrives as a limited beta for users in the United States and can still generate false or contradictory answers. Its `Grok-1` model reaches 73% on MMLU and 63.2% on HumanEval, although it trails GPT-4 in several tests.

xAI has introduced Grok, an artificial intelligence chatbot with access to real-time information through X and a personality designed to answer with humor. The tool arrives as a very early beta and, for now, only a limited group of users in the United States can try it.
A chatbot with information from X
Grok is designed to answer questions on almost any topic, help you come up with new questions and maintain a more irreverent tone than other assistants. According to xAI, its main advantage is that it can access up-to-date information from the X platform.
That could be useful for following an event as it unfolds, summarizing conversations or looking for recent reactions. It also introduces a familiar risk: the assistant could turn incorrect or contradictory information published online into an answer that appears reliable.
xAI warns that Grok can still make up facts. Access to search and real-time information does not eliminate that problem, because the system still generates answers by predicting what text should come next.
What powers Grok
The model behind the chatbot is called Grok-1. xAI says it developed the model in four months after training an earlier prototype called Grok-0, with 33 billion parameters. Parameters are the internal values a model uses to learn language patterns and produce answers.
In the tests published by the company, Grok-1 achieved these results:
- 73% on MMLU, a test with questions from different subjects.
- 63.2% on HumanEval, which measures the ability to complete Python code.
- 62.9% on GSM8K, focused on elementary school math problems.
- 23.9% on MATH, with intermediate and advanced math exercises.
Based on these figures, xAI places Grok above models such as GPT-3.5 and Inflection-1 in its category of training resources. However, GPT-4 performed better in the tests compared: 86.4% on MMLU, 67% on HumanEval and 42.5% on MATH.
These figures are not a definitive ranking of intelligence. Each test uses different instructions, examples and conditions, and xAI itself acknowledges that some of these exams are published online. As a result, it cannot rule out the possibility that the models saw part of their content during training.
A test outside the benchmarks
To reduce that risk, xAI evaluated Grok and other models using Hungary's 2023 national secondary-school math exam. The test was published after the company had collected its training data.
Grok-1 scored 59%, equivalent to a C grade. Claude 2 scored 55% and GPT-4 scored 68%. The result is closer to a real-world situation than a laboratory demonstration, although it still measures only one specific ability.
What changes for you
For now, Grok is not an open service available to everyone. xAI is offering a waitlist for a limited number of users in the United States, who can try the prototype and submit feedback before a wider rollout.
The proposal combines three elements that are already defining AI assistants:
- Answers generated by a language model.
- Searches for recent information on the internet and X.
- A recognizable personality, in this case with humor and a deliberately rebellious tone.
Its connection to X could make Grok useful for current events, but it does not automatically turn its answers into verified facts. You will need to check sensitive data, developing news and any answer that sounds too confident especially carefully.
xAI also explains that it built its infrastructure with Kubernetes, Rust and JAX to train the model using thousands of graphics processors and reduce interruptions during the process. The company is already preparing models with greater capacity and new tools, but it does not present those improvements as available results yet.
Grok marks xAI's entry into the conversational assistant competition, but its starting point is clear: it is a limited beta, with good results on some tests and errors still possible. What matters now is whether access to recent information improves its everyday usefulness without increasing the number of incorrect answers that sound convincing.