<@zlopez:fedora.im>
16:00:14
!startmeeting Infrastructure (2025-09-25)
<@meetbot:fedora.im>
16:00:15
Meeting started at 2025-09-25 16:00:14 UTC
<@meetbot:fedora.im>
16:00:15
The Meeting name is 'Infrastructure (2025-09-25)'
<@zlopez:fedora.im>
16:00:31
!topic Hola y bienvenido
<@zlopez:fedora.im>
16:00:31
!meetingname infrastructure
<@zlopez:fedora.im>
16:00:31
!info Agenda is at: https://board.net/p/fedora-infra
<@zlopez:fedora.im>
16:00:31
!chair @nirik:matrix.scrye.com @zlopez:fedora.im @nb:fedora.im @dtometzki:fedora.im @jnsamyak:matrix.org @james:fedora.im
<@zlopez:fedora.im>
16:00:31
!info About our team: https://docs.fedoraproject.org/en-US/cle/
<@zlopez:fedora.im>
16:00:31
!info Fedora Infra documentation: https://docs.fedoraproject.org/en-US/infra
<@meetbot:fedora.im>
16:00:32
The Meeting Name is now infrastructure
<@zlopez:fedora.im>
16:00:43
Hi everyone
<@nirik:matrix.scrye.com>
16:01:08
morning
<@zlopez:fedora.im>
16:01:12
I'm your host zlopez and this is Jacka... No, it's infrastructure weekly meeting
<@nirik:matrix.scrye.com>
16:03:56
lots of people out on pto I guess... or just busy
<@zlopez:fedora.im>
16:04:11
It's possible
<@zlopez:fedora.im>
16:04:24
I started clearing the chairs as well
<@zlopez:fedora.im>
16:04:51
Waiting for response from some folks
<@zlopez:fedora.im>
16:05:36
Let's continue with it
<@zlopez:fedora.im>
16:05:46
!info Getting Started Guide: https://docs.fedoraproject.org/en-US/infra/gettingstarted/
<@zlopez:fedora.im>
16:05:46
!info This is a place where people who are interested in Fedora Infrastructure can introduce themselves
<@zlopez:fedora.im>
16:05:46
!topic New folks introductions
<@zlopez:fedora.im>
16:05:53
Is anybody new around?
<@zlopez:fedora.im>
16:07:35
It doesn't seem like we have
<@zlopez:fedora.im>
16:07:52
!info magic eight ball says:
<@zlopez:fedora.im>
16:07:52
!info chair 2025-10-02 - nirik
<@zlopez:fedora.im>
16:07:52
!info chair 2025-09-25 - zlopez
<@zlopez:fedora.im>
16:07:52
!topic Next chair
<@zlopez:fedora.im>
16:07:52
!info chair 2025-10-09 - ???
<@zlopez:fedora.im>
16:07:52
!info chair 2025-10-16 - ???
<@zlopez:fedora.im>
16:07:52
!info chair 2025-10-23 - ???
<@zlopez:fedora.im>
16:07:52
!info chair 2025-10-30 - ???
<@zlopez:fedora.im>
16:08:05
So we have a chair for next week
<@nirik:matrix.scrye.com>
16:08:34
yep. I can do it then I think
<@zlopez:fedora.im>
16:08:44
And I'm not here for the next one after, so do we want to leave it empty for now
<@nirik:matrix.scrye.com>
16:08:50
sure
<@zlopez:fedora.im>
16:09:22
In that case let's continue
<@zlopez:fedora.im>
16:09:24
!info Started working on Fedora Infra application list in hackmd https://hackmd.io/@fnF231raRruGKZbPBWKsHQ/B1F1QhRseg [zlopez]
<@zlopez:fedora.im>
16:09:24
!info CLE Infra&Releng EU-hours team has a Monday through Thursday 30 minute meeting going through tickets at 0815 UTC in https://matrix.to/#/#meeting-3:fedoraproject.org
<@zlopez:fedora.im>
16:09:24
topic announcements and information
<@zlopez:fedora.im>
16:09:24
!info CLE Infra&Releng NA-hours team has a Monday through Thursday 30 minute meeting going through tickets at 1900 UTC in https://matrix.to/#/#meeting-3:fedoraproject.org
<@zlopez:fedora.im>
16:09:45
Oh, forgot ! in topic
<@zlopez:fedora.im>
16:09:55
!info https://docs.fedoraproject.org/en-US/infra/day_to_day_fedora/#_the_oncall_role_in_our_team
<@zlopez:fedora.im>
16:09:55
!topic Oncall
<@zlopez:fedora.im>
16:09:55
!info on call from 2025-10-03 to 2025-10-09 - ???
<@zlopez:fedora.im>
16:09:55
!info on call from 2025-10-10 to 2025-10-16 - ???
<@zlopez:fedora.im>
16:09:55
!info on call from 2025-09-26 to 2025-10-02 - zlopez
<@zlopez:fedora.im>
16:09:55
!info on call from 2025-09-19 to 2025-09-25 - nirik
<@zlopez:fedora.im>
16:10:36
I will get to announcements again, I accidentally skipped
<@nirik:matrix.scrye.com>
16:10:49
no biggie
<@zlopez:fedora.im>
16:10:53
I'm probably tired today
<@zlopez:fedora.im>
16:11:45
!oncall
<@zodbot:fedora.im>
16:11:46
The following people are oncall:
<@zodbot:fedora.im>
16:11:46
● @Zlopez:matrix.org (zlopez) Current Time for them: 18:11 (Europe/Prague)
<@zodbot:fedora.im>
16:11:46
If they do not respond, please file a ticket (https://pagure.io/fedora-infrastructure/issues)
<@zodbot:fedora.im>
16:11:46
<@zodbot:fedora.im>
16:11:46
● @gwmngilfen:matrix.org (gwmngilfen) Current Time for them: 17:11 (Europe/London)
<@nirik:matrix.scrye.com>
16:12:06
huh.
<@nirik:matrix.scrye.com>
16:12:12
I could have sworn I changed it.
<@zlopez:fedora.im>
16:12:19
!oncall
<@zodbot:fedora.im>
16:12:20
The following people are oncall:
<@zodbot:fedora.im>
16:12:20
<@zodbot:fedora.im>
16:12:20
● @Zlopez:matrix.org (zlopez) Current Time for them: 18:12 (Europe/Prague)
<@zodbot:fedora.im>
16:12:20
If they do not respond, please file a ticket (https://pagure.io/fedora-infrastructure/issues)
<@zlopez:fedora.im>
16:12:44
Set up, but it seems that Gwmngilfen had oncall this week
<@gwmngilfen:fedora.im>
16:12:56
i did?
<@gwmngilfen:fedora.im>
16:13:10
well, no pings anyway 🙂
<@zlopez:fedora.im>
16:13:10
It looks like that
<@nirik:matrix.scrye.com>
16:13:23
I was supposed to, but somehow I didn't set it... ;(
<@gwmngilfen:fedora.im>
16:13:42
no worries
<@zlopez:fedora.im>
16:13:42
Looking at the history, last change happened 11th September
<@nirik:matrix.scrye.com>
16:13:48
but, I wonder... is oncall really doing much for us? I mean, it seems like a nice idea, but it doesn't get used much anymore...
<@zlopez:fedora.im>
16:14:14
It's true that I didn't saw much oncall pings in the last few weeks
<@gwmngilfen:fedora.im>
16:14:23
i wonder if it would be easier to just add "mods" or "admins" or even "oncall" to our matrix notification wordlist
<@nirik:matrix.scrye.com>
16:14:26
it was used more on irc for some reason... not so much on matrix
<@gwmngilfen:fedora.im>
16:14:34
then if we're around we'll see it
<@nirik:matrix.scrye.com>
16:15:02
well, the idea was that the oncall person could take those inturrupts and everyone else wouldn't have to pay attention/be bothered by it.
<@nirik:matrix.scrye.com>
16:15:50
but given how little it's used, perhaps we should just change it to say 'please describe your issue in #admin:fedoraproject.org and if no response, file a ticket'
<@gwmngilfen:fedora.im>
16:16:12
seems reasonable
<@zlopez:fedora.im>
16:16:15
Maybe that seems like a good way forward
<@nirik:matrix.scrye.com>
16:16:16
(and we would need to update docs for that)
<@nirik:matrix.scrye.com>
16:16:28
(and the bot config)
<@zlopez:fedora.im>
16:16:54
Some default message for the bot
<@nirik:matrix.scrye.com>
16:17:16
or allow it to set a room as oncall instead of a user?
<@gwmngilfen:fedora.im>
16:17:17
one less meeting item is no bad thing either
<@nirik:matrix.scrye.com>
16:17:41
yeah. I can propose it on list/discussion and see if anyone objects?
<@gwmngilfen:fedora.im>
16:17:50
👍️
<@zlopez:fedora.im>
16:18:11
That seems like a good idea
<@nirik:matrix.scrye.com>
16:18:30
added to todo. ✅
<@zlopez:fedora.im>
16:18:43
So let's contineu
<@zlopez:fedora.im>
16:18:55
!info Summary of last week: (from current oncall)
<@zlopez:fedora.im>
16:19:02
I didn't saw anything
<@nirik:matrix.scrye.com>
16:20:03
me either :)
<@gwmngilfen:fedora.im>
16:20:21
🙈
<@zlopez:fedora.im>
16:20:28
Ok, so back to announcements
<@zlopez:fedora.im>
16:20:29
!info Started working on Fedora Infra application list in hackmd https://hackmd.io/@fnF231raRruGKZbPBWKsHQ/B1F1QhRseg [zlopez]
<@zlopez:fedora.im>
16:20:29
!topic announcements and information
<@zlopez:fedora.im>
16:20:29
!info CLE Infra&Releng EU-hours team has a Monday through Thursday 30 minute meeting going through tickets at 0815 UTC in https://matrix.to/#/#meeting-3:fedoraproject.org
<@zlopez:fedora.im>
16:20:29
!info CLE Infra&Releng NA-hours team has a Monday through Thursday 30 minute meeting going through tickets at 1900 UTC in https://matrix.to/#/#meeting-3:fedoraproject.org
<@zlopez:fedora.im>
16:20:44
Anything else to announce?
<@nirik:matrix.scrye.com>
16:21:02
!info anubis is deployed on many of our services, report any issues you see with it.
<@nirik:matrix.scrye.com>
16:21:16
!info mass update/reboot cycle next week
<@gwmngilfen:fedora.im>
16:21:28
i think zabbix will be on that list "soon" - but not this week
<@nirik:matrix.scrye.com>
16:21:33
!info f43 final freeze starting on oct 7th
<@nirik:matrix.scrye.com>
16:23:01
On zabbix... I wonder if the matrix alerts could be threaded (per host?) but thats a RFE not anything big
<@gwmngilfen:fedora.im>
16:23:31
i can investigate but we're using an upstream media type, so ...
<@gwmngilfen:fedora.im>
16:23:48
can always ask / fork, I guess
<@nirik:matrix.scrye.com>
16:23:58
no biggie, just a thought
<@gwmngilfen:fedora.im>
16:24:13
its a good one
<@gwmngilfen:fedora.im>
16:25:20
i cant stay today (will have to go in 5min or so) but if we can schedule ahead I can definitel do a learning sesion on Zabbix one of these weeks
<@zlopez:fedora.im>
16:25:41
Let's continue than
<@zlopez:fedora.im>
16:25:42
!info Go over existing items and fix them
<@zlopez:fedora.im>
16:25:42
!info https://nagios.fedoraproject.org/nagios
<@zlopez:fedora.im>
16:25:42
!topic Monitoring discussion [nirik]
<@zlopez:fedora.im>
16:25:57
I will add that to upcoming learning topics
<@gwmngilfen:fedora.im>
16:26:11
can we make this section Nagios *and* Zabbix? 😛
<@gwmngilfen:fedora.im>
16:26:20
(maybe thats too soon)
<@zlopez:fedora.im>
16:26:48
You should be able to edit agenda as well 🙂
<@nirik:matrix.scrye.com>
16:27:01
so, Gwmngilfen fixed nagios earlier this week... we were not monitoring any of the cloud stuff. ;(
<@nirik:matrix.scrye.com>
16:27:10
and then I fixed up security groups in aws and firewalls.
<@nirik:matrix.scrye.com>
16:27:31
so, we have a few new things:
<@zlopez:fedora.im>
16:27:33
That was a lot of alerts to acknowledge 🙂
<@nirik:matrix.scrye.com>
16:27:53
yeah. ;(
<@nirik:matrix.scrye.com>
16:28:09
bvmhost-p10-01.mgmt.rdu3.fedoraproject.org still is alerting on http I know we fixed it once, but it did not seem to take
<@nirik:matrix.scrye.com>
16:28:47
datanommer01.fedorainfracloud.org is alerting, but I can't seem login to it. :( and actually, perhaps we should just nuke it and the RDS database in favor of the postrest plan?
<@nirik:matrix.scrye.com>
16:29:12
logdetective01.fedorainfracloud.org is alerting and there's a ticket for that team to fix up their firewall
<@nirik:matrix.scrye.com>
16:29:36
certgetter01.rdu3.fedoraproject.org is alerting on http (thats a long standing one we should fix)
<@nirik:matrix.scrye.com>
16:29:59
oci-registry01.rdu3.fedoraproject.org is still alerting on disk space. I guess I should add a few more TB to it. ;(
<@nirik:matrix.scrye.com>
16:30:32
proxy01.rdu3.fedoraproject.org is alerting on the fedoraproject wildcard cert. I plan to renew that this week. (along with several other certs we are not monitoing)
<@nirik:matrix.scrye.com>
16:31:21
Theres a rabbitmq queue alert that I thought I fixed, will make sure it's actually fixed.
<@zlopez:fedora.im>
16:31:48
That is the UNKNOWN one?
<@nirik:matrix.scrye.com>
16:32:03
yeah, that queue was removed...
<@nirik:matrix.scrye.com>
16:32:09
so we shouldn't be monitoring it anymore
<@gwmngilfen:fedora.im>
16:32:21
Zabbix side:
<@gwmngilfen:fedora.im>
16:32:21
- I pushed a bunch of threshold updates to quieten down some of the noisy hosts
<@gwmngilfen:fedora.im>
16:32:21
- i've fixed up a bunch of openvpn certs today that have restored agent connectivity for vmhost-x86-cc03/5/6 and ibiblio02
<@gwmngilfen:fedora.im>
16:32:54
theres some disk space issues in there too which need consideration (or new thresholds)
<@gwmngilfen:fedora.im>
16:33:37
log01 and ipa01 are both warning about / being over 80% full
<@nirik:matrix.scrye.com>
16:33:59
I guess nagios has a higher threshold there?
<@gwmngilfen:fedora.im>
16:34:09
probably. i can look and mirror it
<@nirik:matrix.scrye.com>
16:34:29
ipa01 is probibly something causing undue logs.
<@nirik:matrix.scrye.com>
16:34:34
log01 ditto. ;)
<@gwmngilfen:fedora.im>
16:34:42
these are new-ish alerts so something must be using disk
<@nirik:matrix.scrye.com>
16:35:35
log01 was almost full due to a crashlooping poddler... the compressed gigantic logs probibly are still taking up a lot of room
<@nirik:matrix.scrye.com>
16:35:48
it was around 100GB or so of logs
<@gwmngilfen:fedora.im>
16:35:53
ah, fair enough, let me ack that for now then
<@nirik:matrix.scrye.com>
16:35:53
(uncompressed)
<@zlopez:fedora.im>
16:36:01
That is fixed now
<@gwmngilfen:fedora.im>
16:36:25
oh an people01 is high-cpu again, looks like cgit is stuck again
<@gwmngilfen:fedora.im>
16:36:28
oh and people01 is high-cpu again, looks like cgit is stuck again
<@nirik:matrix.scrye.com>
16:36:35
yeah, but the old logs (now rotated and compressed) might be taking up space... will have to look
<@gwmngilfen:fedora.im>
16:37:15
staging Zabbix looks fine though 😛 ok, gotta run, I'll be on mobile if you need my input
<@zlopez:fedora.im>
16:37:37
On ipa01 I can see 4.1 GB on httpd logs
<@zlopez:fedora.im>
16:37:59
There are plenty of uncompressed ones
<@nirik:matrix.scrye.com>
16:38:40
anyhow, I think thats it for monitoring... we can file tickets/solve the issues out of meeting
<@zlopez:fedora.im>
16:39:07
Ok, what do we want to do next?
<@nirik:matrix.scrye.com>
16:39:27
I'd say backlog, but it's only 2 of us, so perhaps we just end early.
<@zlopez:fedora.im>
16:40:12
I'm +1 for ending early
<@zlopez:fedora.im>
16:40:36
!topic Open Floor
<@zlopez:fedora.im>
16:41:06
Let's switch to open floor for few minutes and I will end the meeting after that
<@zlopez:fedora.im>
16:41:22
So anybody who wants to say something have 5 minutes for now
<@nirik:matrix.scrye.com>
16:42:35
I'd tell a UDP joke here, but I am not sure you all would get it?
<@zlopez:fedora.im>
16:47:13
Thanks everyone for coming today
<@zlopez:fedora.im>
16:47:16
!endmeeting