mirror of
https://github.com/FEX-Emu/FEX.git
synced 2026-10-08 16:00:19 +02:00
Compare commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
d3399a261b | ||
|
|
d2437e6a21 | ||
|
|
95dd6ceba8 | ||
|
|
1a0d135201 | ||
|
|
f453e1523e | ||
|
|
622b0bfbc9 | ||
|
|
ad52514b97 | ||
|
|
02a218c6e3 | ||
|
|
2d617ad173 | ||
|
|
0d06e3e47d | ||
|
|
78aee4d96e | ||
|
|
ba04da87e5 | ||
|
|
2e6b08cbcb | ||
|
|
472a373861 | ||
|
|
a451420911 | ||
|
|
2e84f21c18 | ||
|
|
fb7167c2d2 | ||
|
|
d884eb9287 | ||
|
|
8b9b1a90e4 | ||
|
|
e2d4010b59 | ||
|
|
babde31bf0 | ||
|
|
c282239077 | ||
|
|
8d28a441ab | ||
|
|
5821054d91 | ||
|
|
4626145374 | ||
|
|
a786d3621d | ||
|
|
1393dc2a5b | ||
|
|
c4604465ba | ||
|
|
cffae9cb0f | ||
|
|
cf24d3c33f | ||
|
|
672e885e40 | ||
|
|
7d05610da7 | ||
|
|
a843ecf4c8 | ||
|
|
f4ff1b0688 | ||
|
|
2b4cec8385 | ||
|
|
8ab4ab29f8 | ||
|
|
cc0509c0f3 | ||
|
|
ebfa65fedc | ||
|
|
a34ae24b3f | ||
|
|
58ea76eb24 | ||
|
|
e9a17b19c5 | ||
|
|
ce8d111453 | ||
|
|
47fd73f6cf | ||
|
|
76f3391ebc | ||
|
|
be6ff52709 | ||
|
|
e99e252188 | ||
|
|
98b980f7e3 | ||
|
|
f2f90eeb82 | ||
|
|
f267fd2250 | ||
|
|
500ad34769 | ||
|
|
1700d54012 | ||
|
|
70d8a10484 | ||
|
|
9e94784e26 | ||
|
|
4060f4018e | ||
|
|
739ac0f18f | ||
|
|
98d62a7eb1 | ||
|
|
aba7a3a830 | ||
|
|
9027d1eee7 | ||
|
|
4e5da4946d | ||
|
|
a70e3e42b2 | ||
|
|
09f476924f | ||
|
|
230e3245fd | ||
|
|
8de876daf2 | ||
|
|
53b1d155cc | ||
|
|
b0eb63ab9a | ||
|
|
2e3242682d | ||
|
|
ad4d4c9e67 | ||
|
|
3250d4e405 | ||
|
|
196a0531e0 | ||
|
|
e61cb5b2c3 | ||
|
|
f9b53c6b51 | ||
|
|
58e949e148 | ||
|
|
dad47b7bda | ||
|
|
e519bf5978 | ||
|
|
fc50e52157 | ||
|
|
7669df0e16 | ||
|
|
4d56fec5f1 | ||
|
|
8181552b16 | ||
|
|
c6c147daf6 | ||
|
|
975069825e | ||
|
|
5133f480d1 | ||
|
|
ce4b252e5c | ||
|
|
031d56de35 | ||
|
|
3cdaf6736b | ||
|
|
b5e696b3cb | ||
|
|
43aef377d7 | ||
|
|
add0e7a8db | ||
|
|
52e541d453 | ||
|
|
a031a49546 | ||
|
|
4d821b8dd8 | ||
|
|
f277025c9a | ||
|
|
ba28e6f82e | ||
|
|
3a89df9bed | ||
|
|
f6a0866fbb | ||
|
|
756fa2ecc5 | ||
|
|
cf834aa6da | ||
|
|
d2324f4a93 | ||
|
|
6226c7f4f3 | ||
|
|
991ecd558e | ||
|
|
a4fa3a460e | ||
|
|
77ba708933 | ||
|
|
662d50a966 | ||
|
|
5472d1cc04 | ||
|
|
d1d41f5645 | ||
|
|
94fd100fc7 | ||
|
|
b9ff36b5d9 | ||
|
|
cd5a809ec9 | ||
|
|
045a8efbeb | ||
|
|
54a1f7d833 | ||
|
|
1b496cda8f | ||
|
|
a5b24bfe4c | ||
|
|
46676ca376 | ||
|
|
7d939a3b3d | ||
|
|
a515061465 | ||
|
|
1c24d63f73 | ||
|
|
7e10dba5e2 | ||
|
|
e2d73014f1 | ||
|
|
53aa30596e | ||
|
|
122ae5b710 | ||
|
|
45c27b2965 | ||
|
|
832b247fc1 | ||
|
|
0e8b53d566 | ||
|
|
d03d69273b | ||
|
|
efa05ba19d | ||
|
|
5da205d91a | ||
|
|
41923bac99 | ||
|
|
c6148f6bf1 | ||
|
|
77aaa9af4d | ||
|
|
00cf8d530c | ||
|
|
98aa58e9f5 | ||
|
|
6911917819 | ||
|
|
a8255aa475 | ||
|
|
48e7aae38f | ||
|
|
7069643ae6 | ||
|
|
34272fc134 | ||
|
|
1d41002dfe | ||
|
|
563bf342d5 | ||
|
|
efd5fabb95 | ||
|
|
c1da525110 | ||
|
|
5ce6c88a88 | ||
|
|
64cce7c6fa | ||
|
|
4544e5b51f | ||
|
|
eb3e314946 | ||
|
|
8b65c3de10 | ||
|
|
c2beb27a9d | ||
|
|
05fdec9e72 | ||
|
|
e8e3c95349 | ||
|
|
a87fa3f246 | ||
|
|
b31ad523f5 | ||
|
|
34bce540ff | ||
|
|
8ea38e1d80 | ||
|
|
ce591a9541 | ||
|
|
a48c65cd65 | ||
|
|
c283f80f48 | ||
|
|
d6bf276b5a | ||
|
|
96a51650b1 | ||
|
|
f35a9c74a2 | ||
|
|
e2457943f5 | ||
|
|
a05644172a | ||
|
|
cc168ce0fb | ||
|
|
76bd22d279 | ||
|
|
18574f3cf1 | ||
|
|
665215ab47 | ||
|
|
6009f36403 | ||
|
|
2580efda0d | ||
|
|
3a310b8815 | ||
|
|
7ff96227c0 | ||
|
|
3e8d78051c | ||
|
|
bd24ebc96a | ||
|
|
6d3745b8f1 | ||
|
|
d0f0b975be | ||
|
|
ff2e6ed59f | ||
|
|
99b2018d0e | ||
|
|
f0d9c8c10a | ||
|
|
dce1b24c00 | ||
|
|
b47e981932 | ||
|
|
dc44eb4caf | ||
|
|
dfda6733f0 | ||
|
|
21c6986dc7 | ||
|
|
635720fe12 | ||
|
|
8c751d7423 | ||
|
|
702ecf7637 | ||
|
|
0595f1e044 | ||
|
|
cebb032bd3 | ||
|
|
8e32763ada | ||
|
|
7532337231 | ||
|
|
4a66d4570e | ||
|
|
6f5e99d47d | ||
|
|
9ee9f5bddd | ||
|
|
ddb9f6d3ad | ||
|
|
d29139d88a | ||
|
|
317575ba99 | ||
|
|
d4f2638a2e | ||
|
|
b67d9be227 | ||
|
|
d52add8fad | ||
|
|
aa9159d25c | ||
|
|
94c777259e | ||
|
|
c9f8fa5662 | ||
|
|
64ee6b119e | ||
|
|
d2ec9a8936 | ||
|
|
2a927453f7 | ||
|
|
c19d489c9a | ||
|
|
6012eb051b | ||
|
|
3974746473 | ||
|
|
e1bcdcf387 | ||
|
|
fd5fbddae9 | ||
|
|
8ff72beddb | ||
|
|
cba5f7877b | ||
|
|
9d7e9fd9fc | ||
|
|
082a0baff3 | ||
|
|
3a4914315b | ||
|
|
448b5a338a | ||
|
|
9b68617fa8 | ||
|
|
4c9890d7f8 | ||
|
|
b2db04f5d7 | ||
|
|
be8ff9ccb9 | ||
|
|
9c531d97b0 | ||
|
|
055d8d75a2 | ||
|
|
9fcf79ce0e | ||
|
|
6edf4619d4 | ||
|
|
8f769ce5a3 | ||
|
|
96ac71750a | ||
|
|
d0852cf1bb | ||
|
|
f5fea8af96 | ||
|
|
d52a1da501 | ||
|
|
abdcaa7c86 | ||
|
|
ad122cf463 | ||
|
|
b58a57d225 | ||
|
|
28d679de98 | ||
|
|
d1dd055e6a | ||
|
|
3045578da4 | ||
|
|
9566dda73e | ||
|
|
a0ced2b685 | ||
|
|
df232f567b | ||
|
|
2a6d6a9d13 | ||
|
|
cd03932bd1 | ||
|
|
2e5fa1ef1b | ||
|
|
0c6c4cd532 | ||
|
|
25f8a87429 | ||
|
|
edf1a7970d | ||
|
|
fac9972bad | ||
|
|
9ecb960f3a | ||
|
|
3d26e23891 | ||
|
|
7bbbd95775 | ||
|
|
bb308899b9 | ||
|
|
e95c8d703c | ||
|
|
903d6a742e | ||
|
|
424218e327 | ||
|
|
17dc03d414 | ||
|
|
baf699c6e1 | ||
|
|
1431af1ff5 | ||
|
|
775a41b903 | ||
|
|
2da1e90dd5 | ||
|
|
e614340c0c | ||
|
|
3c293b9aed | ||
|
|
283c2861c9 | ||
|
|
757dc95116 | ||
|
|
6192250b8a | ||
|
|
f489135b1d | ||
|
|
4d00a52761 | ||
|
|
6941a59223 | ||
|
|
3f232e631e | ||
|
|
6e3643c3ef | ||
|
|
d7348c8aff | ||
|
|
e7bdb8679d | ||
|
|
c28824f94d | ||
|
|
664d766b45 | ||
|
|
fce694ed92 | ||
|
|
96aafb4f07 | ||
|
|
a474f86ea8 | ||
|
|
dbaf95a8f3 | ||
|
|
e67df96ad9 | ||
|
|
56de94578d | ||
|
|
06fc2f5ef0 | ||
|
|
b3ba315cbd | ||
|
|
e5a531e683 | ||
|
|
e2de57bd04 | ||
|
|
4eebca93e3 | ||
|
|
3919ec9692 | ||
|
|
02aeb0ac1a | ||
|
|
206544ad09 | ||
|
|
3854cd2b2f | ||
|
|
b2eb8aaf66 | ||
|
|
acbd920c9a | ||
|
|
db0bdd48e5 | ||
|
|
da21ee3cda | ||
|
|
d2baef2b36 | ||
|
|
df96bc83cc | ||
|
|
ec03831a21 | ||
|
|
9ca821316a | ||
|
|
025a060337 | ||
|
|
371d6f0730 | ||
|
|
643bc10d52 | ||
|
|
8fb801069f | ||
|
|
542ed8b6ad | ||
|
|
053620f4f5 | ||
|
|
88b01a0ca9 | ||
|
|
197140498b | ||
|
|
9acd325aa4 | ||
|
|
2483329ef6 | ||
|
|
f6b58b4219 | ||
|
|
6c6d86f761 | ||
|
|
359221b379 | ||
|
|
87fe1d672e | ||
|
|
9257221b3b | ||
|
|
f9b38a1de7 | ||
|
|
67e1ac0442 | ||
|
|
c57e9e008f | ||
|
|
b34c23fe3d | ||
|
|
29f644235d | ||
|
|
01da5972fc | ||
|
|
643e964edd | ||
|
|
30e3d795da | ||
|
|
89b05a2ea4 | ||
|
|
2e009be27c | ||
|
|
32150cf7b5 | ||
|
|
bf812aae8f | ||
|
|
9a71443005 | ||
|
|
ee165249bc | ||
|
|
7c7d767195 | ||
|
|
af8cfb79e5 | ||
|
|
27c8bf3021 | ||
|
|
b0a09b31bb | ||
|
|
13ebfb1a49 | ||
|
|
f863b30951 | ||
|
|
1ce27a5e6b | ||
|
|
933d622860 | ||
|
|
5d67223236 | ||
|
|
825d2c948c | ||
|
|
29390b439a | ||
|
|
799c17eb90 | ||
|
|
5fb84866e0 | ||
|
|
4965344ef5 | ||
|
|
46ca53ad0d | ||
|
|
61ff1b3584 | ||
|
|
7c0c5de4bd | ||
|
|
8d134b8df8 | ||
|
|
a9bacc1b6b | ||
|
|
e4ff3dac86 | ||
|
|
9443b18076 | ||
|
|
bb4e81aa19 | ||
|
|
a9a9f6782a | ||
|
|
2fa6c3c918 | ||
|
|
4bd84eb523 | ||
|
|
e2073dcd30 | ||
|
|
fd72669c7e | ||
|
|
81c144697b | ||
|
|
aecf180dfe | ||
|
|
10fa4a4f20 | ||
|
|
534732564b | ||
|
|
6a314bc9cd | ||
|
|
1d4356b97e | ||
|
|
a11566012d | ||
|
|
8d929027c8 | ||
|
|
1d1ed012d8 | ||
|
|
b092b7a937 | ||
|
|
d133fa6dc1 | ||
|
|
afa7de969e | ||
|
|
41b6a89ffd | ||
|
|
9aa82ec5bf | ||
|
|
b17a2e9f96 | ||
|
|
af4e9ceeed | ||
|
|
9c62c41f5f | ||
|
|
184c9d21bb | ||
|
|
9744d8de99 | ||
|
|
f3e6ecb2c3 | ||
|
|
aa0f2c3975 | ||
|
|
4dc8648d81 | ||
|
|
7a703e1176 | ||
|
|
02df1a2924 | ||
|
|
efac7efc97 | ||
|
|
d99b4a80c8 | ||
|
|
2a76744d30 | ||
|
|
843b2d1969 | ||
|
|
b275c96889 | ||
|
|
033b1ce449 | ||
|
|
99a43283be | ||
|
|
55bfd6394b | ||
|
|
14bfe6016e | ||
|
|
0d4ad70875 | ||
|
|
a8bf3859ea | ||
|
|
aa7dcffcea | ||
|
|
be1a5cea8e | ||
|
|
402ea84aa0 | ||
|
|
19a7b06b91 | ||
|
|
96bd643e5b | ||
|
|
6b9293979c | ||
|
|
7d5cee4384 | ||
|
|
c0bab70161 | ||
|
|
32f5a28433 | ||
|
|
ce30179ed1 | ||
|
|
a515b707f3 | ||
|
|
9ab0fa01bd | ||
|
|
c3bffa2929 | ||
|
|
951fee361f | ||
|
|
8c4860b9a7 | ||
|
|
ee221e6a8c | ||
|
|
f5625093bb | ||
|
|
abfd974d70 | ||
|
|
97966930e9 | ||
|
|
a52a2e3ae4 | ||
|
|
c49b30f105 | ||
|
|
b0b4ad2083 | ||
|
|
ee4bee4fef | ||
|
|
c3a0f5a2f6 | ||
|
|
0413a6bf68 | ||
|
|
7bd036d1ae | ||
|
|
112c49a348 | ||
|
|
80878ae611 | ||
|
|
85a69be5b6 | ||
|
|
8dbfd1635a | ||
|
|
8b5ca303e3 | ||
|
|
f90d2aeb6d | ||
|
|
cd4b520f72 | ||
|
|
20d5a26a72 | ||
|
|
9346116485 | ||
|
|
bb8336fcad | ||
|
|
ee96d60983 | ||
|
|
6052b335dc | ||
|
|
ab0a6bbe9f | ||
|
|
9dd6d8ed94 | ||
|
|
3b5d0e3e27 | ||
|
|
37e13cf073 | ||
|
|
9b1b9c26cc | ||
|
|
f7f3024b92 | ||
|
|
11ec71a4ce | ||
|
|
665491adf8 | ||
|
|
55391ccbc0 | ||
|
|
7790d7a0b7 | ||
|
|
f3d8c2cbac | ||
|
|
1226069b4c | ||
|
|
80687c8d2d | ||
|
|
920fe60492 | ||
|
|
8c6ce2cb3b | ||
|
|
61f30d004c | ||
|
|
95919a1ddf | ||
|
|
35ec54f920 | ||
|
|
32e8a56093 | ||
|
|
136f1d0a0b | ||
|
|
0c042d1e85 | ||
|
|
ad13442be4 | ||
|
|
d6b9252760 | ||
|
|
22222ebaf5 | ||
|
|
734258e23b | ||
|
|
74916b3757 | ||
|
|
c5359264a3 | ||
|
|
9d0ff7929e | ||
|
|
d3eed27d17 | ||
|
|
bc1669b163 | ||
|
|
83e417b2c6 | ||
|
|
cb00d9171f | ||
|
|
cf77f2ae5d | ||
|
|
273d086a7b | ||
|
|
94d9cf54bc | ||
|
|
3089e0e6de | ||
|
|
3c088fb414 | ||
|
|
676c9a9be6 | ||
|
|
314fea36b4 | ||
|
|
3b7d30d26a | ||
|
|
a8d32b9a2f | ||
|
|
24cb02f4ff | ||
|
|
725d0e187a | ||
|
|
d8603cb9bd | ||
|
|
04c2cb5feb | ||
|
|
32f2decb24 | ||
|
|
6954ebe3a0 | ||
|
|
2e40da3d6b | ||
|
|
559772bb03 | ||
|
|
2bbcf72e27 | ||
|
|
aa3a92aa60 | ||
|
|
5497240a25 | ||
|
|
79c609d0f5 | ||
|
|
b33788e765 | ||
|
|
8340012466 | ||
|
|
0adcc779cf | ||
|
|
9d0718fbc4 | ||
|
|
53567a6526 | ||
|
|
90fb5f038b | ||
|
|
063b1eb936 | ||
|
|
9f243c8f7b | ||
|
|
2dc600f283 | ||
|
|
a01402d502 | ||
|
|
c90036aeea | ||
|
|
dfb751eea0 | ||
|
|
6e0f5eccb3 | ||
|
|
579fb42458 | ||
|
|
f4b487352c | ||
|
|
4448f84f29 | ||
|
|
9e1e602e09 | ||
|
|
ca70e387ec | ||
|
|
9a483107e3 | ||
|
|
3bac767866 | ||
|
|
7b4e48480b | ||
|
|
101bba4808 | ||
|
|
a31c3c1c15 | ||
|
|
1a18e392f8 | ||
|
|
4c7595c68a | ||
|
|
1a467f0ebd | ||
|
|
06e7360f4c | ||
|
|
ec3b72e17e | ||
|
|
259e1b75a4 | ||
|
|
f9642cba7a | ||
|
|
c01c415030 | ||
|
|
50e56358c3 | ||
|
|
465dbc260f | ||
|
|
9f291f3adb | ||
|
|
f2acc3da4c | ||
|
|
f8f165d96d | ||
|
|
baef95992c | ||
|
|
97e18c8469 | ||
|
|
6e04f7368b | ||
|
|
55a835ebb8 | ||
|
|
85776c2537 | ||
|
|
e3ec25d9db | ||
|
|
ebfcc1e835 | ||
|
|
769b2c2a46 | ||
|
|
3b2100307e | ||
|
|
28cc179214 | ||
|
|
3c0f243a2d | ||
|
|
bbf1563f80 | ||
|
|
ed6b1011f9 | ||
|
|
e3e7f0279c | ||
|
|
a10f984b1c | ||
|
|
663f3d8b5a | ||
|
|
b1f7be2f6c | ||
|
|
b83cbcb33c | ||
|
|
ac1a096bae | ||
|
|
048c8ded88 | ||
|
|
948938bf4b | ||
|
|
bb064c7334 | ||
|
|
2d3d49b900 | ||
|
|
ea7096ed5b | ||
|
|
7a0f6c0a80 | ||
|
|
e4ee35a925 | ||
|
|
d3ab9bdef6 | ||
|
|
926eefc86c | ||
|
|
3eb7a5b998 | ||
|
|
58614ff131 | ||
|
|
bf3a09e5e3 | ||
|
|
83c536c47f | ||
|
|
7b39e57e72 | ||
|
|
5bedf32666 | ||
|
|
1eb7be9870 | ||
|
|
93e4288c57 | ||
|
|
9ca4868833 | ||
|
|
c5f8ea58e9 | ||
|
|
5bee17bee1 | ||
|
|
efe7c54374 | ||
|
|
9e1840e974 | ||
|
|
1f40590f9a | ||
|
|
64a3bc235d | ||
|
|
7d9af246ea | ||
|
|
6d3471bcaa | ||
|
|
a8714dbd49 | ||
|
|
512312fa06 | ||
|
|
010028e381 | ||
|
|
f27f1871e4 | ||
|
|
3a7aa83ab1 | ||
|
|
bcc136c3b9 | ||
|
|
3da31830d1 | ||
|
|
d19b57a52e | ||
|
|
ef6d640a8c | ||
|
|
2cae2f2462 | ||
|
|
1fde5d7fca | ||
|
|
10de2f83ac | ||
|
|
9d86e11a47 | ||
|
|
4d503d3155 | ||
|
|
7e663b91df | ||
|
|
47242dc190 | ||
|
|
e13c8e3295 | ||
|
|
3c3ba62c10 | ||
|
|
34fe56dfb2 | ||
|
|
ecf8cde5e0 | ||
|
|
55284aad7e | ||
|
|
3afc35f7b4 | ||
|
|
1058428a51 | ||
|
|
a2fc51fc7b | ||
|
|
1848629ba5 | ||
|
|
3399577330 | ||
|
|
18bfc8afd0 | ||
|
|
76b023ed3e | ||
|
|
b91b0e9d65 | ||
|
|
74489a4177 | ||
|
|
55d1d6bcd4 | ||
|
|
cd249e2c3a | ||
|
|
61cd835754 | ||
|
|
bd24364c1b | ||
|
|
ab516d7b79 | ||
|
|
d25ed4b0bf | ||
|
|
c521d2b48d | ||
|
|
9ed8165405 | ||
|
|
5099b2b5dc | ||
|
|
d372552593 | ||
|
|
729e32ccc2 | ||
|
|
170204d6f1 | ||
|
|
f7bfecd3f1 | ||
|
|
5f0427c253 | ||
|
|
472860a840 | ||
|
|
eddb7d12cc | ||
|
|
789a9f19c0 | ||
|
|
f70aafb211 | ||
|
|
7519af2819 |
No files matched your search
+1
-1
@@ -7,7 +7,7 @@ AlignConsecutiveAssignments: None
|
||||
AlignConsecutiveBitFields: Consecutive
|
||||
AlignConsecutiveDeclarations: None
|
||||
AlignConsecutiveMacros: None
|
||||
AlignEscapedNewlines: DontAlign
|
||||
AlignEscapedNewlines: Left
|
||||
AlignOperands: Align
|
||||
AlignTrailingComments: true
|
||||
AllowAllParametersOfDeclarationOnNextLine: false
|
||||
|
||||
@@ -8,7 +8,5 @@ FEXCore/Source/Common/SoftFloat-3e/*
|
||||
Source/Common/cpp-optparse/*
|
||||
|
||||
# Files with human-indented tables for readability - don't mess with these
|
||||
FEXCore/Source/Interface/Core/X86Tables/X87Tables.cpp
|
||||
FEXCore/Source/Interface/Core/X86Tables/XOPTables.cpp
|
||||
FEXCore/Source/Interface/Core/X86Tables/*
|
||||
|
||||
@@ -13,7 +13,6 @@ env:
|
||||
BUILD_TYPE: Release
|
||||
CC: clang
|
||||
CXX: clang++
|
||||
FEX_ENABLEAVX: 1
|
||||
|
||||
jobs:
|
||||
build_plus_test:
|
||||
@@ -64,7 +63,7 @@ jobs:
|
||||
# Note the current convention is to use the -S and -B options here to specify source
|
||||
# and build directories, but this is only available with CMake 3.13 and higher.
|
||||
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
|
||||
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DBUILD_FEX_LINUX_TESTS=True -DBUILD_THUNKS=True -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
|
||||
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_FEX_LINUX_TESTS=True -DBUILD_THUNKS=True -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
|
||||
|
||||
- name: Build
|
||||
working-directory: ${{runner.workspace}}/build
|
||||
|
||||
@@ -20,7 +20,6 @@ env:
|
||||
BUILD_TYPE: Release
|
||||
CC: clang
|
||||
CXX: clang++
|
||||
FEX_ENABLEAVX: 1
|
||||
|
||||
jobs:
|
||||
glibc_fault_test:
|
||||
@@ -71,7 +70,7 @@ jobs:
|
||||
# Note the current convention is to use the -S and -B options here to specify source
|
||||
# and build directories, but this is only available with CMake 3.13 and higher.
|
||||
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
|
||||
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DBUILD_FEX_LINUX_TESTS=True -DENABLE_GLIBC_ALLOCATOR_HOOK_FAULT=True -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
|
||||
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_FEX_LINUX_TESTS=True -DENABLE_GLIBC_ALLOCATOR_HOOK_FAULT=True -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
|
||||
|
||||
- name: Build
|
||||
working-directory: ${{runner.workspace}}/build
|
||||
|
||||
@@ -13,7 +13,6 @@ env:
|
||||
BUILD_TYPE: Release
|
||||
CC: clang
|
||||
CXX: clang++
|
||||
FEX_ENABLEAVX: 1
|
||||
|
||||
jobs:
|
||||
hostrunner_tests:
|
||||
@@ -64,7 +63,7 @@ jobs:
|
||||
# Note the current convention is to use the -S and -B options here to specify source
|
||||
# and build directories, but this is only available with CMake 3.13 and higher.
|
||||
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
|
||||
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True
|
||||
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True
|
||||
|
||||
- name: Build
|
||||
working-directory: ${{runner.workspace}}/build
|
||||
|
||||
@@ -13,7 +13,6 @@ env:
|
||||
BUILD_TYPE: Release
|
||||
CC: clang
|
||||
CXX: clang++
|
||||
FEX_ENABLEAVX: 1
|
||||
|
||||
jobs:
|
||||
instcountci_tests:
|
||||
@@ -74,7 +73,7 @@ jobs:
|
||||
# Note the current convention is to use the -S and -B options here to specify source
|
||||
# and build directories, but this is only available with CMake 3.13 and higher.
|
||||
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
|
||||
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_VIXL_SIMULATOR=$VIXL_SIM_ENABLED -DENABLE_VIXL_DISASSEMBLER=True -DENABLE_LTO=False -DENABLE_ASSERTIONS=True
|
||||
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_VIXL_SIMULATOR=$VIXL_SIM_ENABLED -DENABLE_VIXL_DISASSEMBLER=True -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True
|
||||
|
||||
- name: Build
|
||||
working-directory: ${{runner.workspace}}/build
|
||||
|
||||
@@ -10,7 +10,6 @@ on:
|
||||
|
||||
env:
|
||||
BUILD_TYPE: Debug
|
||||
FEX_ENABLEAVX: 1
|
||||
|
||||
jobs:
|
||||
mingw_build:
|
||||
@@ -74,7 +73,7 @@ jobs:
|
||||
# Note the current convention is to use the -S and -B options here to specify source
|
||||
# and build directories, but this is only available with CMake 3.13 and higher.
|
||||
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
|
||||
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/toolchain_mingw.cmake -DMINGW_TRIPLE=$MINGW_TRIPLE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DBUILD_TESTS=False -DENABLE_JEMALLOC=False -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
|
||||
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -DCMAKE_TOOLCHAIN_FILE=$GITHUB_WORKSPACE/toolchain_mingw.cmake -DMINGW_TRIPLE=$MINGW_TRIPLE -G Ninja -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True -DBUILD_TESTS=False -DENABLE_JEMALLOC=False -DENABLE_JEMALLOC_GLIBC_ALLOC=False -DCMAKE_INSTALL_PREFIX=${{runner.workspace}}/build/install
|
||||
|
||||
- name: Build
|
||||
working-directory: ${{runner.workspace}}/build
|
||||
|
||||
@@ -20,6 +20,7 @@ jobs:
|
||||
|
||||
- name: Checkout through merge base
|
||||
uses: rmacklin/fetch-through-merge-base@v0
|
||||
timeout-minutes: 3
|
||||
with:
|
||||
base_ref: ${{ github.event.pull_request.base.ref }}
|
||||
head_ref: ${{ github.event.pull_request.head.sha }}
|
||||
|
||||
@@ -13,7 +13,6 @@ env:
|
||||
BUILD_TYPE: Release
|
||||
CC: clang
|
||||
CXX: clang++
|
||||
FEX_ENABLEAVX: 1
|
||||
|
||||
jobs:
|
||||
vixl_simulator:
|
||||
@@ -65,7 +64,7 @@ jobs:
|
||||
# Note the current convention is to use the -S and -B options here to specify source
|
||||
# and build directories, but this is only available with CMake 3.13 and higher.
|
||||
# The CMake binaries on the Github Actions machines are (as of this writing) 3.12
|
||||
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_VIXL_SIMULATOR=True -DENABLE_VIXL_DISASSEMBLER=True -DENABLE_LTO=False -DENABLE_ASSERTIONS=True
|
||||
run: cmake $GITHUB_WORKSPACE -DCMAKE_BUILD_TYPE=$BUILD_TYPE -G Ninja -DENABLE_VIXL_SIMULATOR=True -DENABLE_VIXL_DISASSEMBLER=True -DENABLE_LTO=False -DENABLE_ASSERTIONS=True -DENABLE_X86_HOST_DEBUG=True
|
||||
|
||||
- name: Build
|
||||
working-directory: ${{runner.workspace}}/build
|
||||
|
||||
+24
-2
@@ -43,6 +43,14 @@ if (NOT CONTAINS_MINGW EQUAL -1)
|
||||
set (ENABLE_JEMALLOC FALSE)
|
||||
endif()
|
||||
|
||||
if (NOT MINGW_BUILD)
|
||||
message (STATUS "Clang version ${CMAKE_CXX_COMPILER_VERSION}")
|
||||
set (CLANG_MINIMUM_VERSION 12.0)
|
||||
if (CMAKE_CXX_COMPILER_VERSION VERSION_LESS ${CLANG_MINIMUM_VERSION})
|
||||
message (FATAL_ERROR "Clang version too old for FEX. Need at least ${CLANG_MINIMUM_VERSION} but has ${CMAKE_CXX_COMPILER_VERSION}")
|
||||
endif()
|
||||
endif()
|
||||
|
||||
if (ENABLE_FEXCORE_PROFILER)
|
||||
add_definitions(-DENABLE_FEXCORE_PROFILER=1)
|
||||
string(TOUPPER "${FEXCORE_PROFILER_BACKEND}" FEXCORE_PROFILER_BACKEND)
|
||||
@@ -110,6 +118,12 @@ else()
|
||||
endif()
|
||||
|
||||
if (CMAKE_SYSTEM_PROCESSOR MATCHES "x86_64")
|
||||
option(ENABLE_X86_HOST_DEBUG "Enables compiling on x86_64 host" FALSE)
|
||||
if (NOT ENABLE_X86_HOST_DEBUG)
|
||||
message(FATAL_ERROR
|
||||
" FEX-Emu doesn't support compiling for x86-64 hosts!"
|
||||
" This is /only/ a supported configuration for FEX CI and nothing else!")
|
||||
endif()
|
||||
set(_M_X86_64 1)
|
||||
add_definitions(-D_M_X86_64=1)
|
||||
set (CMAKE_CXX_FLAGS "${CMAKE_CXX_FLAGS} -mcx16")
|
||||
@@ -235,8 +249,6 @@ add_definitions(-Wno-trigraphs)
|
||||
add_definitions(-DGLOBAL_DATA_DIRECTORY="${DATA_DIRECTORY}/")
|
||||
|
||||
if (BUILD_TESTS)
|
||||
option(CATCH_BUILD_STATIC_LIBRARY "" ON)
|
||||
set(CATCH_BUILD_STATIC_LIBRARY ON)
|
||||
add_subdirectory(External/Catch2/)
|
||||
|
||||
# Pull in catch_discover_tests definition
|
||||
@@ -356,9 +368,19 @@ if (BUILD_TESTS)
|
||||
include(CTest)
|
||||
enable_testing()
|
||||
message(STATUS "Unit tests are enabled")
|
||||
|
||||
set (TEST_JOB_COUNT "" CACHE STRING "Override number of parallel jobs to use while running tests")
|
||||
if (TEST_JOB_COUNT)
|
||||
message(STATUS "Running tests with ${TEST_JOB_COUNT} jobs")
|
||||
endif()
|
||||
if (CMAKE_VERSION VERSION_LESS "3.29")
|
||||
execute_process(COMMAND "nproc" OUTPUT_STRIP_TRAILING_WHITESPACE OUTPUT_VARIABLE TEST_JOB_COUNT)
|
||||
endif()
|
||||
set(TEST_JOB_FLAG "-j${TEST_JOB_COUNT}")
|
||||
endif()
|
||||
|
||||
add_subdirectory(FEXHeaderUtils/)
|
||||
add_subdirectory(CodeEmitter/)
|
||||
add_subdirectory(FEXCore/)
|
||||
|
||||
# Binfmt_misc files must be installed prior to Source/ installs
|
||||
|
||||
@@ -0,0 +1,2 @@
|
||||
add_library(CodeEmitter INTERFACE)
|
||||
target_include_directories(CodeEmitter INTERFACE .)
|
||||
File diff suppressed because it is too large.
Load diff
+1302
-1303
File diff suppressed because it is too large.
Load diff
+32
-32
@@ -8,11 +8,11 @@ public:
|
||||
public:
|
||||
// Conditional branch immediate
|
||||
///< Branch conditional
|
||||
void b(FEXCore::ARMEmitter::Condition Cond, uint32_t Imm) {
|
||||
void b(ARMEmitter::Condition Cond, uint32_t Imm) {
|
||||
constexpr uint32_t Op = 0b0101'010 << 25;
|
||||
Branch_Conditional(Op, 0, 0, Cond, Imm);
|
||||
}
|
||||
void b(FEXCore::ARMEmitter::Condition Cond, BackwardLabel const* Label) {
|
||||
void b(ARMEmitter::Condition Cond, BackwardLabel const* Label) {
|
||||
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
|
||||
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
|
||||
constexpr uint32_t Op = 0b0101'010 << 25;
|
||||
@@ -20,13 +20,13 @@ public:
|
||||
}
|
||||
template<typename LabelType>
|
||||
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
|
||||
void b(FEXCore::ARMEmitter::Condition Cond, LabelType *Label) {
|
||||
void b(ARMEmitter::Condition Cond, LabelType *Label) {
|
||||
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::BC });
|
||||
constexpr uint32_t Op = 0b0101'010 << 25;
|
||||
Branch_Conditional(Op, 0, 0, Cond, 0);
|
||||
}
|
||||
|
||||
void b(FEXCore::ARMEmitter::Condition Cond, BiDirectionalLabel *Label) {
|
||||
void b(ARMEmitter::Condition Cond, BiDirectionalLabel *Label) {
|
||||
if (Label->Backward.Location) {
|
||||
b(Cond, &Label->Backward);
|
||||
}
|
||||
@@ -36,11 +36,11 @@ public:
|
||||
}
|
||||
|
||||
///< Branch consistent conditional
|
||||
void bc(FEXCore::ARMEmitter::Condition Cond, uint32_t Imm) {
|
||||
void bc(ARMEmitter::Condition Cond, uint32_t Imm) {
|
||||
constexpr uint32_t Op = 0b0101'010 << 25;
|
||||
Branch_Conditional(Op, 0, 1, Cond, Imm);
|
||||
}
|
||||
void bc(FEXCore::ARMEmitter::Condition Cond, BackwardLabel const* Label) {
|
||||
void bc(ARMEmitter::Condition Cond, BackwardLabel const* Label) {
|
||||
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
|
||||
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
|
||||
constexpr uint32_t Op = 0b0101'010 << 25;
|
||||
@@ -49,13 +49,13 @@ public:
|
||||
|
||||
template<typename LabelType>
|
||||
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
|
||||
void bc(FEXCore::ARMEmitter::Condition Cond, LabelType *Label) {
|
||||
void bc(ARMEmitter::Condition Cond, LabelType *Label) {
|
||||
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::BC });
|
||||
constexpr uint32_t Op = 0b0101'010 << 25;
|
||||
Branch_Conditional(Op, 0, 1, Cond, 0);
|
||||
}
|
||||
|
||||
void bc(FEXCore::ARMEmitter::Condition Cond, BiDirectionalLabel *Label) {
|
||||
void bc(ARMEmitter::Condition Cond, BiDirectionalLabel *Label) {
|
||||
if (Label->Backward.Location) {
|
||||
bc(Cond, &Label->Backward);
|
||||
}
|
||||
@@ -65,7 +65,7 @@ public:
|
||||
}
|
||||
|
||||
// Unconditional branch register
|
||||
void br(FEXCore::ARMEmitter::Register rn) {
|
||||
void br(ARMEmitter::Register rn) {
|
||||
constexpr uint32_t Op = 0b1101011 << 25 |
|
||||
0b0'000 << 21 | // opc
|
||||
0b1'1111 << 16 | // op2
|
||||
@@ -74,7 +74,7 @@ public:
|
||||
|
||||
UnconditionalBranch(Op, rn);
|
||||
}
|
||||
void blr(FEXCore::ARMEmitter::Register rn) {
|
||||
void blr(ARMEmitter::Register rn) {
|
||||
constexpr uint32_t Op = 0b1101011 << 25 |
|
||||
0b0'001 << 21 | // opc
|
||||
0b1'1111 << 16 | // op2
|
||||
@@ -83,7 +83,7 @@ public:
|
||||
|
||||
UnconditionalBranch(Op, rn);
|
||||
}
|
||||
void ret(FEXCore::ARMEmitter::Register rn = FEXCore::ARMEmitter::Reg::r30) {
|
||||
void ret(ARMEmitter::Register rn = ARMEmitter::Reg::r30) {
|
||||
constexpr uint32_t Op = 0b1101011 << 25 |
|
||||
0b0'010 << 21 | // opc
|
||||
0b1'1111 << 16 | // op2
|
||||
@@ -156,13 +156,13 @@ public:
|
||||
}
|
||||
|
||||
// Compare and branch
|
||||
void cbz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, uint32_t Imm) {
|
||||
void cbz(ARMEmitter::Size s, ARMEmitter::Register rt, uint32_t Imm) {
|
||||
constexpr uint32_t Op = 0b0011'0100 << 24;
|
||||
|
||||
CompareAndBranch(Op, s, rt, Imm);
|
||||
}
|
||||
|
||||
void cbz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, BackwardLabel const* Label) {
|
||||
void cbz(ARMEmitter::Size s, ARMEmitter::Register rt, BackwardLabel const* Label) {
|
||||
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
|
||||
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
|
||||
|
||||
@@ -173,7 +173,7 @@ public:
|
||||
|
||||
template<typename LabelType>
|
||||
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
|
||||
void cbz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, LabelType *Label) {
|
||||
void cbz(ARMEmitter::Size s, ARMEmitter::Register rt, LabelType *Label) {
|
||||
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::BC });
|
||||
|
||||
constexpr uint32_t Op = 0b0011'0100 << 24;
|
||||
@@ -181,7 +181,7 @@ public:
|
||||
CompareAndBranch(Op, s, rt, 0);
|
||||
}
|
||||
|
||||
void cbz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, BiDirectionalLabel *Label) {
|
||||
void cbz(ARMEmitter::Size s, ARMEmitter::Register rt, BiDirectionalLabel *Label) {
|
||||
if (Label->Backward.Location) {
|
||||
cbz(s, rt, &Label->Backward);
|
||||
}
|
||||
@@ -190,13 +190,13 @@ public:
|
||||
}
|
||||
}
|
||||
|
||||
void cbnz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, uint32_t Imm) {
|
||||
void cbnz(ARMEmitter::Size s, ARMEmitter::Register rt, uint32_t Imm) {
|
||||
constexpr uint32_t Op = 0b0011'0101 << 24;
|
||||
|
||||
CompareAndBranch(Op, s, rt, Imm);
|
||||
}
|
||||
|
||||
void cbnz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, BackwardLabel const* Label) {
|
||||
void cbnz(ARMEmitter::Size s, ARMEmitter::Register rt, BackwardLabel const* Label) {
|
||||
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
|
||||
LOGMAN_THROW_A_FMT(Imm >= -1048576 && Imm <= 1048575 && ((Imm & 0b11) == 0), "Unscaled offset too large");
|
||||
|
||||
@@ -207,7 +207,7 @@ public:
|
||||
|
||||
template<typename LabelType>
|
||||
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
|
||||
void cbnz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, LabelType *Label) {
|
||||
void cbnz(ARMEmitter::Size s, ARMEmitter::Register rt, LabelType *Label) {
|
||||
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::BC });
|
||||
|
||||
constexpr uint32_t Op = 0b0011'0101 << 24;
|
||||
@@ -215,7 +215,7 @@ public:
|
||||
CompareAndBranch(Op, s, rt, 0);
|
||||
}
|
||||
|
||||
void cbnz(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, BiDirectionalLabel *Label) {
|
||||
void cbnz(ARMEmitter::Size s, ARMEmitter::Register rt, BiDirectionalLabel *Label) {
|
||||
if (Label->Backward.Location) {
|
||||
cbnz(s, rt, &Label->Backward);
|
||||
}
|
||||
@@ -225,12 +225,12 @@ public:
|
||||
}
|
||||
|
||||
// Test and branch immediate
|
||||
void tbz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, uint32_t Imm) {
|
||||
void tbz(ARMEmitter::Register rt, uint32_t Bit, uint32_t Imm) {
|
||||
constexpr uint32_t Op = 0b0011'0110 << 24;
|
||||
|
||||
TestAndBranch(Op, rt, Bit, Imm);
|
||||
}
|
||||
void tbz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, BackwardLabel const* Label) {
|
||||
void tbz(ARMEmitter::Register rt, uint32_t Bit, BackwardLabel const* Label) {
|
||||
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
|
||||
LOGMAN_THROW_A_FMT(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0), "Unscaled offset too large");
|
||||
|
||||
@@ -241,7 +241,7 @@ public:
|
||||
|
||||
template<typename LabelType>
|
||||
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
|
||||
void tbz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, LabelType *Label) {
|
||||
void tbz(ARMEmitter::Register rt, uint32_t Bit, LabelType *Label) {
|
||||
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::TEST_BRANCH });
|
||||
|
||||
constexpr uint32_t Op = 0b0011'0110 << 24;
|
||||
@@ -249,7 +249,7 @@ public:
|
||||
TestAndBranch(Op, rt, Bit, 0);
|
||||
}
|
||||
|
||||
void tbz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, BiDirectionalLabel *Label) {
|
||||
void tbz(ARMEmitter::Register rt, uint32_t Bit, BiDirectionalLabel *Label) {
|
||||
if (Label->Backward.Location) {
|
||||
tbz(rt, Bit, &Label->Backward);
|
||||
}
|
||||
@@ -258,12 +258,12 @@ public:
|
||||
}
|
||||
}
|
||||
|
||||
void tbnz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, uint32_t Imm) {
|
||||
void tbnz(ARMEmitter::Register rt, uint32_t Bit, uint32_t Imm) {
|
||||
constexpr uint32_t Op = 0b0011'0111 << 24;
|
||||
|
||||
TestAndBranch(Op, rt, Bit, Imm);
|
||||
}
|
||||
void tbnz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, BackwardLabel const* Label) {
|
||||
void tbnz(ARMEmitter::Register rt, uint32_t Bit, BackwardLabel const* Label) {
|
||||
int32_t Imm = static_cast<int32_t>(Label->Location - GetCursorAddress<uint8_t*>());
|
||||
LOGMAN_THROW_A_FMT(Imm >= -32768 && Imm <= 32764 && ((Imm & 0b11) == 0), "Unscaled offset too large");
|
||||
|
||||
@@ -274,14 +274,14 @@ public:
|
||||
|
||||
template<typename LabelType>
|
||||
requires (std::is_same_v<LabelType, ForwardLabel> || std::is_same_v<LabelType, SingleUseForwardLabel>)
|
||||
void tbnz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, LabelType *Label) {
|
||||
void tbnz(ARMEmitter::Register rt, uint32_t Bit, LabelType *Label) {
|
||||
AddLocationToLabel(Label, SingleUseForwardLabel{ .Location = GetCursorAddress<uint8_t*>(), .Type = SingleUseForwardLabel::InstType::TEST_BRANCH });
|
||||
constexpr uint32_t Op = 0b0011'0111 << 24;
|
||||
|
||||
TestAndBranch(Op, rt, Bit, 0);
|
||||
}
|
||||
|
||||
void tbnz(FEXCore::ARMEmitter::Register rt, uint32_t Bit, BiDirectionalLabel *Label) {
|
||||
void tbnz(ARMEmitter::Register rt, uint32_t Bit, BiDirectionalLabel *Label) {
|
||||
if (Label->Backward.Location) {
|
||||
tbnz(rt, Bit, &Label->Backward);
|
||||
}
|
||||
@@ -292,7 +292,7 @@ public:
|
||||
|
||||
private:
|
||||
// Conditional branch immediate
|
||||
void Branch_Conditional(uint32_t Op, uint32_t Op1, uint32_t Op0, FEXCore::ARMEmitter::Condition Cond, uint32_t Imm) {
|
||||
void Branch_Conditional(uint32_t Op, uint32_t Op1, uint32_t Op0, ARMEmitter::Condition Cond, uint32_t Imm) {
|
||||
uint32_t Instr = Op;
|
||||
|
||||
Instr |= Op1 << 24;
|
||||
@@ -304,7 +304,7 @@ private:
|
||||
}
|
||||
|
||||
// Unconditional branch register
|
||||
void UnconditionalBranch(uint32_t Op, FEXCore::ARMEmitter::Register rn) {
|
||||
void UnconditionalBranch(uint32_t Op, ARMEmitter::Register rn) {
|
||||
uint32_t Instr = Op;
|
||||
Instr |= Encode_rn(rn);
|
||||
dc32(Instr);
|
||||
@@ -318,8 +318,8 @@ private:
|
||||
}
|
||||
|
||||
// Compare and branch
|
||||
void CompareAndBranch(uint32_t Op, FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register rt, uint32_t Imm) {
|
||||
const uint32_t SF = s == FEXCore::ARMEmitter::Size::i64Bit ? (1U << 31) : 0;
|
||||
void CompareAndBranch(uint32_t Op, ARMEmitter::Size s, ARMEmitter::Register rt, uint32_t Imm) {
|
||||
const uint32_t SF = s == ARMEmitter::Size::i64Bit ? (1U << 31) : 0;
|
||||
|
||||
uint32_t Instr = Op;
|
||||
|
||||
@@ -330,7 +330,7 @@ private:
|
||||
}
|
||||
|
||||
// Test and branch - immediate
|
||||
void TestAndBranch(uint32_t Op, FEXCore::ARMEmitter::Register rt, uint32_t Bit, uint32_t Imm) {
|
||||
void TestAndBranch(uint32_t Op, ARMEmitter::Register rt, uint32_t Bit, uint32_t Imm) {
|
||||
uint32_t Instr = Op;
|
||||
|
||||
Instr |= (Bit >> 5) << 31;
|
||||
+2
-2
@@ -4,7 +4,7 @@
|
||||
#include <cstdint>
|
||||
#include <cstring>
|
||||
|
||||
namespace FEXCore::ARMEmitter {
|
||||
namespace ARMEmitter {
|
||||
class Buffer {
|
||||
public:
|
||||
Buffer() {
|
||||
@@ -103,4 +103,4 @@ protected:
|
||||
uint8_t* CurrentOffset;
|
||||
uint64_t Size;
|
||||
};
|
||||
} // namespace FEXCore::ARMEmitter
|
||||
} // namespace ARMEmitter
|
||||
+14
-15
@@ -1,9 +1,6 @@
|
||||
// SPDX-License-Identifier: MIT
|
||||
#pragma once
|
||||
|
||||
#include "Interface/Core/ArchHelpers/CodeEmitter/Buffer.h"
|
||||
#include "Interface/Core/ArchHelpers/CodeEmitter/Registers.h"
|
||||
|
||||
#include <FEXCore/Utils/CompilerDefs.h>
|
||||
#include <FEXCore/Utils/EnumUtils.h>
|
||||
#include <FEXCore/Utils/LogManager.h>
|
||||
@@ -11,8 +8,8 @@
|
||||
#include <FEXCore/fextl/vector.h>
|
||||
|
||||
#include <FEXHeaderUtils/BitUtils.h>
|
||||
|
||||
#include <aarch64/assembler-aarch64.h>
|
||||
#include <CodeEmitter/Buffer.h>
|
||||
#include <CodeEmitter/Registers.h>
|
||||
|
||||
#include <array>
|
||||
#include <cstdint>
|
||||
@@ -56,7 +53,7 @@
|
||||
* it easier to select the correct load-store instruction. Mostly because these are a nightmare selecting
|
||||
* the right instruction.
|
||||
*/
|
||||
namespace FEXCore::ARMEmitter {
|
||||
namespace ARMEmitter {
|
||||
/*
|
||||
* This `Size` enum is used for most ALU operations.
|
||||
* These follow the AArch64 encoding style in most cases.
|
||||
@@ -611,7 +608,7 @@ constexpr bool AreVectorsSequential(T first, const Args&... args) {
|
||||
// Choices:
|
||||
// - Size of ops passed as an argument rather than template to let the compiler use csel instead of branching.
|
||||
// - Registers are unsized so they can be passed in a GPR and not need conversion operations
|
||||
class Emitter : public FEXCore::ARMEmitter::Buffer {
|
||||
class Emitter : public ARMEmitter::Buffer {
|
||||
public:
|
||||
Emitter() = default;
|
||||
|
||||
@@ -759,15 +756,17 @@ public:
|
||||
Bind<false>(&Label->Forward);
|
||||
}
|
||||
|
||||
#include <CodeEmitter/VixlUtils.inl>
|
||||
|
||||
public:
|
||||
// TODO: Implement SME when it matters.
|
||||
#include "Interface/Core/ArchHelpers/CodeEmitter/ALUOps.inl"
|
||||
#include "Interface/Core/ArchHelpers/CodeEmitter/BranchOps.inl"
|
||||
#include "Interface/Core/ArchHelpers/CodeEmitter/LoadstoreOps.inl"
|
||||
#include "Interface/Core/ArchHelpers/CodeEmitter/SystemOps.inl"
|
||||
#include "Interface/Core/ArchHelpers/CodeEmitter/ScalarOps.inl"
|
||||
#include "Interface/Core/ArchHelpers/CodeEmitter/ASIMDOps.inl"
|
||||
#include "Interface/Core/ArchHelpers/CodeEmitter/SVEOps.inl"
|
||||
#include <CodeEmitter/ALUOps.inl>
|
||||
#include <CodeEmitter/BranchOps.inl>
|
||||
#include <CodeEmitter/LoadstoreOps.inl>
|
||||
#include <CodeEmitter/SystemOps.inl>
|
||||
#include <CodeEmitter/ScalarOps.inl>
|
||||
#include <CodeEmitter/ASIMDOps.inl>
|
||||
#include <CodeEmitter/SVEOps.inl>
|
||||
|
||||
private:
|
||||
template<typename T>
|
||||
@@ -829,4 +828,4 @@ private:
|
||||
return FEXCore::ToUnderlying(Reg);
|
||||
}
|
||||
};
|
||||
} // namespace FEXCore::ARMEmitter
|
||||
} // namespace ARMEmitter
|
||||
+481
-487
File diff suppressed because it is too large.
Load diff
+2
-2
@@ -6,7 +6,7 @@
|
||||
#include <compare>
|
||||
#include <cstdint>
|
||||
|
||||
namespace FEXCore::ARMEmitter {
|
||||
namespace ARMEmitter {
|
||||
class WRegister;
|
||||
class XRegister;
|
||||
|
||||
@@ -1024,4 +1024,4 @@ enum class OpType : uint32_t {
|
||||
Destructive = 0,
|
||||
Constructive,
|
||||
};
|
||||
} // namespace FEXCore::ARMEmitter
|
||||
} // namespace ARMEmitter
|
||||
+42
-26
@@ -1506,13 +1506,14 @@ public:
|
||||
}
|
||||
|
||||
// SVE broadcast floating-point immediate (unpredicated)
|
||||
void fdup(FEXCore::ARMEmitter::SubRegSize size, FEXCore::ARMEmitter::ZRegister zd, float Value) {
|
||||
LOGMAN_THROW_AA_FMT(size == FEXCore::ARMEmitter::SubRegSize::i16Bit ||
|
||||
size == FEXCore::ARMEmitter::SubRegSize::i32Bit ||
|
||||
size == FEXCore::ARMEmitter::SubRegSize::i64Bit, "Unsupported fmov size");
|
||||
void fdup(ARMEmitter::SubRegSize size, ARMEmitter::ZRegister zd, float Value) {
|
||||
LOGMAN_THROW_AA_FMT(size == ARMEmitter::SubRegSize::i16Bit ||
|
||||
size == ARMEmitter::SubRegSize::i32Bit ||
|
||||
size == ARMEmitter::SubRegSize::i64Bit, "Unsupported fmov size");
|
||||
uint32_t Imm{};
|
||||
if (size == SubRegSize::i16Bit) {
|
||||
Imm = FP16ToImm8(vixl::Float16(Value));
|
||||
LOGMAN_MSG_A_FMT("Unsupported");
|
||||
FEX_UNREACHABLE;
|
||||
} else if (size == SubRegSize::i32Bit) {
|
||||
Imm = FP32ToImm8(Value);
|
||||
} else if (size == SubRegSize::i64Bit) {
|
||||
@@ -1521,7 +1522,7 @@ public:
|
||||
|
||||
SVEBroadcastFloatImmUnpredicated(0b00, 0, Imm, size, zd);
|
||||
}
|
||||
void fmov(FEXCore::ARMEmitter::SubRegSize size, FEXCore::ARMEmitter::ZRegister zd, float Value) {
|
||||
void fmov(ARMEmitter::SubRegSize size, ARMEmitter::ZRegister zd, float Value) {
|
||||
fdup(size, zd, Value);
|
||||
}
|
||||
|
||||
@@ -3320,7 +3321,18 @@ public:
|
||||
|
||||
// SVE Memory - Contiguous Store with Immediate Offset
|
||||
// SVE contiguous non-temporal store (scalar plus immediate)
|
||||
// XXX:
|
||||
void stnt1b(ZRegister zt, PRegister pg, Register rn, int32_t Imm = 0) {
|
||||
SVEContiguousNontemporalStore(0b00, zt, pg, rn, Imm);
|
||||
}
|
||||
void stnt1h(ZRegister zt, PRegister pg, Register rn, int32_t Imm = 0) {
|
||||
SVEContiguousNontemporalStore(0b01, zt, pg, rn, Imm);
|
||||
}
|
||||
void stnt1w(ZRegister zt, PRegister pg, Register rn, int32_t Imm = 0) {
|
||||
SVEContiguousNontemporalStore(0b10, zt, pg, rn, Imm);
|
||||
}
|
||||
void stnt1d(ZRegister zt, PRegister pg, Register rn, int32_t Imm = 0) {
|
||||
SVEContiguousNontemporalStore(0b11, zt, pg, rn, Imm);
|
||||
}
|
||||
|
||||
// SVE store multiple structures (scalar plus immediate)
|
||||
void st2b(ZRegister zt1, ZRegister zt2, PRegister pg, Register rn, int32_t Imm = 0) {
|
||||
@@ -3513,7 +3525,8 @@ private:
|
||||
size == SubRegSize::i64Bit, "Unsupported fcpy/fmov size");
|
||||
uint32_t imm{};
|
||||
if (size == SubRegSize::i16Bit) {
|
||||
imm = FP16ToImm8(vixl::Float16(value));
|
||||
LOGMAN_MSG_A_FMT("Unsupported");
|
||||
FEX_UNREACHABLE;
|
||||
} else if (size == SubRegSize::i32Bit) {
|
||||
imm = FP32ToImm8(value);
|
||||
} else if (size == SubRegSize::i64Bit) {
|
||||
@@ -3717,7 +3730,7 @@ private:
|
||||
|
||||
// SVE bitwise logical operations (predicated)
|
||||
void SVEBitwiseLogicalPredicated(uint32_t opc, SubRegSize size, PRegister pg, ZRegister zdn, ZRegister zm, ZRegister zd) {
|
||||
LOGMAN_THROW_AA_FMT(size != FEXCore::ARMEmitter::SubRegSize::i128Bit, "Can't use 128-bit size");
|
||||
LOGMAN_THROW_AA_FMT(size != ARMEmitter::SubRegSize::i128Bit, "Can't use 128-bit size");
|
||||
LOGMAN_THROW_A_FMT(zd == zdn, "zd needs to equal zdn");
|
||||
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
|
||||
|
||||
@@ -4479,6 +4492,22 @@ private:
|
||||
dc32(Instr);
|
||||
}
|
||||
|
||||
// SVE contiguous non-temporal store (scalar plus immediate)
|
||||
void SVEContiguousNontemporalStore(uint32_t msz, ZRegister zt, PRegister pg, Register rn, int32_t imm) {
|
||||
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
|
||||
LOGMAN_THROW_AA_FMT(imm >= -8 && imm <= 7,
|
||||
"Invalid loadstore offset ({}). Must be between [-8, 7]", imm);
|
||||
|
||||
const auto imm4 = static_cast<uint32_t>(imm) & 0xF;
|
||||
uint32_t Instr = 0b1110'0100'0001'0000'1110'0000'0000'0000;
|
||||
Instr |= msz << 23;
|
||||
Instr |= imm4 << 16;
|
||||
Instr |= pg.Idx() << 10;
|
||||
Instr |= Encode_rn(rn);
|
||||
Instr |= zt.Idx();
|
||||
dc32(Instr);
|
||||
}
|
||||
|
||||
void SVEContiguousLoadImm(bool is_store, uint32_t dtype, int32_t imm, PRegister pg, Register rn, ZRegister zt) {
|
||||
LOGMAN_THROW_A_FMT(pg <= PReg::p7, "Can only use p0-p7 as a governing predicate");
|
||||
LOGMAN_THROW_AA_FMT(imm >= -8 && imm <= 7,
|
||||
@@ -4743,7 +4772,7 @@ private:
|
||||
dc32(Instr);
|
||||
}
|
||||
|
||||
void SVEPermuteVector(uint32_t op0, FEXCore::ARMEmitter::ZRegister zd, FEXCore::ARMEmitter::ZRegister zm, uint32_t Imm) {
|
||||
void SVEPermuteVector(uint32_t op0, ARMEmitter::ZRegister zd, ARMEmitter::ZRegister zm, uint32_t Imm) {
|
||||
constexpr uint32_t Op = 0b0000'0101'0010'0000'000 << 13;
|
||||
uint32_t Instr = Op;
|
||||
|
||||
@@ -5228,15 +5257,14 @@ private:
|
||||
|
||||
// Alias that returns the equivalently sized unsigned type for a floating-point type T.
|
||||
template <typename T>
|
||||
requires(std::is_same_v<T, float> || std::is_same_v<T, double> || std::is_same_v<T, vixl::Float16>)
|
||||
using FloatToEquivalentUInt = std::conditional_t<std::is_same_v<T, vixl::Float16>, uint16_t,
|
||||
std::conditional_t<std::is_same_v<T, float>, uint32_t, uint64_t>>;
|
||||
requires(std::is_same_v<T, float> || std::is_same_v<T, double>)
|
||||
using FloatToEquivalentUInt = std::conditional_t<std::is_same_v<T, float>, uint32_t, uint64_t>;
|
||||
|
||||
// Determines if a floating-point value is capable of being converted
|
||||
// into an 8-bit immediate. See pseudocode definition of VFPExpandImm
|
||||
// in ARM A-profile reference manual for a general overview of how this was derived.
|
||||
template <typename T>
|
||||
requires(std::is_same_v<T, float> || std::is_same_v<T, double> || std::is_same_v<T, vixl::Float16>)
|
||||
requires(std::is_same_v<T, float> || std::is_same_v<T, double>)
|
||||
[[nodiscard, maybe_unused]] static bool IsValidFPValueForImm8(T value) {
|
||||
const uint64_t bits = FEXCore::BitCast<FloatToEquivalentUInt<T>>(value);
|
||||
const uint64_t datasize_idx = FEXCore::ilog2(sizeof(T)) - 1;
|
||||
@@ -5277,18 +5305,6 @@ private:
|
||||
return true;
|
||||
}
|
||||
|
||||
static uint32_t FP16ToImm8(vixl::Float16 value) {
|
||||
LOGMAN_THROW_A_FMT(IsValidFPValueForImm8(value),
|
||||
"Value cannot be encoded into an 8-bit immediate");
|
||||
|
||||
const uint32_t bits = vixl::Float16ToRawbits(value);
|
||||
const uint32_t sign = (bits & 0x8000) >> 8;
|
||||
const uint32_t expb2 = (bits & 0x2000) >> 7;
|
||||
const uint32_t b5_to_0 = (bits >> 6) & 0x3F;
|
||||
|
||||
return sign | expb2 | b5_to_0;
|
||||
}
|
||||
|
||||
static uint32_t FP32ToImm8(float value) {
|
||||
LOGMAN_THROW_A_FMT(IsValidFPValueForImm8(value),
|
||||
"Value ({}) cannot be encoded into an 8-bit immediate", value);
|
||||
+9
-9
@@ -33,7 +33,7 @@ public:
|
||||
ASIMDScalarCopy(Op, 1, imm5, 0b0000, rd, rn);
|
||||
}
|
||||
|
||||
void mov(FEXCore::ARMEmitter::ScalarRegSize size, FEXCore::ARMEmitter::VRegister rd, FEXCore::ARMEmitter::VRegister rn, uint32_t Index) {
|
||||
void mov(ARMEmitter::ScalarRegSize size, ARMEmitter::VRegister rd, ARMEmitter::VRegister rn, uint32_t Index) {
|
||||
dup(size, rd, rn, Index);
|
||||
}
|
||||
|
||||
@@ -1052,21 +1052,21 @@ public:
|
||||
}
|
||||
|
||||
// Floating-point immediate
|
||||
void fmov(FEXCore::ARMEmitter::ScalarRegSize size, FEXCore::ARMEmitter::VRegister rd, float Value) {
|
||||
void fmov(ARMEmitter::ScalarRegSize size, ARMEmitter::VRegister rd, float Value) {
|
||||
uint32_t M = 0;
|
||||
uint32_t S = 0;
|
||||
uint32_t ptype;
|
||||
uint32_t imm8;
|
||||
uint32_t imm5 = 0b0'0000;
|
||||
if (size == FEXCore::ARMEmitter::ScalarRegSize::i16Bit) {
|
||||
ptype = 0b11;
|
||||
imm8 = FP16ToImm8(vixl::Float16(Value));
|
||||
if (size == ARMEmitter::ScalarRegSize::i16Bit) {
|
||||
LOGMAN_MSG_A_FMT("Unsupported");
|
||||
FEX_UNREACHABLE;
|
||||
}
|
||||
else if (size == FEXCore::ARMEmitter::ScalarRegSize::i32Bit) {
|
||||
else if (size == ARMEmitter::ScalarRegSize::i32Bit) {
|
||||
ptype = 0b00;
|
||||
imm8 = FP32ToImm8(Value);
|
||||
}
|
||||
else if (size == FEXCore::ARMEmitter::ScalarRegSize::i64Bit) {
|
||||
else if (size == ARMEmitter::ScalarRegSize::i64Bit) {
|
||||
ptype = 0b01;
|
||||
imm8 = FP64ToImm8(Value);
|
||||
}
|
||||
@@ -1077,7 +1077,7 @@ public:
|
||||
FloatScalarImmediate(M, S, ptype, imm8, imm5, rd);
|
||||
}
|
||||
|
||||
void FloatScalarImmediate(uint32_t M, uint32_t S, uint32_t ptype, uint32_t imm8, uint32_t imm5, FEXCore::ARMEmitter::VRegister rd) {
|
||||
void FloatScalarImmediate(uint32_t M, uint32_t S, uint32_t ptype, uint32_t imm8, uint32_t imm5, ARMEmitter::VRegister rd) {
|
||||
constexpr uint32_t Op = 0b0001'1110'0010'0000'0001'00 << 10;
|
||||
uint32_t Instr = Op;
|
||||
|
||||
@@ -1286,7 +1286,7 @@ public:
|
||||
|
||||
private:
|
||||
// Advanced SIMD scalar copy
|
||||
void ASIMDScalarCopy(uint32_t Op, uint32_t Q, uint32_t imm5, uint32_t imm4, FEXCore::ARMEmitter::VRegister rd, FEXCore::ARMEmitter::VRegister rn) {
|
||||
void ASIMDScalarCopy(uint32_t Op, uint32_t Q, uint32_t imm5, uint32_t imm4, ARMEmitter::VRegister rd, ARMEmitter::VRegister rn) {
|
||||
uint32_t Instr = Op;
|
||||
|
||||
Instr |= Q << 30;
|
||||
+26
-26
@@ -11,7 +11,7 @@ public:
|
||||
// TODO: AT
|
||||
// TODO: CFP
|
||||
// TODO: CPP
|
||||
void dc(FEXCore::ARMEmitter::DataCacheOperation DCOp, FEXCore::ARMEmitter::Register rt) {
|
||||
void dc(ARMEmitter::DataCacheOperation DCOp, ARMEmitter::Register rt) {
|
||||
constexpr uint32_t Op = 0b1101'0101'0000'1000'0111 << 12;
|
||||
SystemInstruction(Op, 0, FEXCore::ToUnderlying(DCOp), rt);
|
||||
}
|
||||
@@ -48,67 +48,67 @@ public:
|
||||
ExceptionGeneration(0b101, 0b000, 0b11, Imm);
|
||||
}
|
||||
// System instructions with register argument
|
||||
void wfet(FEXCore::ARMEmitter::Register rt) {
|
||||
void wfet(ARMEmitter::Register rt) {
|
||||
SystemInstructionWithReg(0b0000, 0b000, rt);
|
||||
}
|
||||
void wfit(FEXCore::ARMEmitter::Register rt) {
|
||||
void wfit(ARMEmitter::Register rt) {
|
||||
SystemInstructionWithReg(0b0000, 0b001, rt);
|
||||
}
|
||||
|
||||
// Hints
|
||||
void nop() {
|
||||
Hint(FEXCore::ARMEmitter::HintRegister::NOP);
|
||||
Hint(ARMEmitter::HintRegister::NOP);
|
||||
}
|
||||
void yield() {
|
||||
Hint(FEXCore::ARMEmitter::HintRegister::YIELD);
|
||||
Hint(ARMEmitter::HintRegister::YIELD);
|
||||
}
|
||||
void wfe() {
|
||||
Hint(FEXCore::ARMEmitter::HintRegister::WFE);
|
||||
Hint(ARMEmitter::HintRegister::WFE);
|
||||
}
|
||||
void wfi() {
|
||||
Hint(FEXCore::ARMEmitter::HintRegister::WFI);
|
||||
Hint(ARMEmitter::HintRegister::WFI);
|
||||
}
|
||||
void sev() {
|
||||
Hint(FEXCore::ARMEmitter::HintRegister::SEV);
|
||||
Hint(ARMEmitter::HintRegister::SEV);
|
||||
}
|
||||
void sevl() {
|
||||
Hint(FEXCore::ARMEmitter::HintRegister::SEVL);
|
||||
Hint(ARMEmitter::HintRegister::SEVL);
|
||||
}
|
||||
void dgh() {
|
||||
Hint(FEXCore::ARMEmitter::HintRegister::DGH);
|
||||
Hint(ARMEmitter::HintRegister::DGH);
|
||||
}
|
||||
void csdb() {
|
||||
Hint(FEXCore::ARMEmitter::HintRegister::CSDB);
|
||||
Hint(ARMEmitter::HintRegister::CSDB);
|
||||
}
|
||||
|
||||
// Barriers
|
||||
void clrex(uint32_t imm = 15) {
|
||||
LOGMAN_THROW_AA_FMT(imm < 16, "Immediate out of range");
|
||||
Barrier(FEXCore::ARMEmitter::BarrierRegister::CLREX, imm);
|
||||
Barrier(ARMEmitter::BarrierRegister::CLREX, imm);
|
||||
}
|
||||
void dsb(FEXCore::ARMEmitter::BarrierScope Scope) {
|
||||
Barrier(FEXCore::ARMEmitter::BarrierRegister::DSB, FEXCore::ToUnderlying(Scope));
|
||||
void dsb(ARMEmitter::BarrierScope Scope) {
|
||||
Barrier(ARMEmitter::BarrierRegister::DSB, FEXCore::ToUnderlying(Scope));
|
||||
}
|
||||
void dmb(FEXCore::ARMEmitter::BarrierScope Scope) {
|
||||
Barrier(FEXCore::ARMEmitter::BarrierRegister::DMB, FEXCore::ToUnderlying(Scope));
|
||||
void dmb(ARMEmitter::BarrierScope Scope) {
|
||||
Barrier(ARMEmitter::BarrierRegister::DMB, FEXCore::ToUnderlying(Scope));
|
||||
}
|
||||
void isb() {
|
||||
Barrier(FEXCore::ARMEmitter::BarrierRegister::ISB, FEXCore::ToUnderlying(FEXCore::ARMEmitter::BarrierScope::SY));
|
||||
Barrier(ARMEmitter::BarrierRegister::ISB, FEXCore::ToUnderlying(ARMEmitter::BarrierScope::SY));
|
||||
}
|
||||
void sb() {
|
||||
Barrier(FEXCore::ARMEmitter::BarrierRegister::SB, 0);
|
||||
Barrier(ARMEmitter::BarrierRegister::SB, 0);
|
||||
}
|
||||
void tcommit() {
|
||||
Barrier(FEXCore::ARMEmitter::BarrierRegister::TCOMMIT, 0);
|
||||
Barrier(ARMEmitter::BarrierRegister::TCOMMIT, 0);
|
||||
}
|
||||
|
||||
// System register move
|
||||
void msr(FEXCore::ARMEmitter::SystemRegister reg, FEXCore::ARMEmitter::Register rt) {
|
||||
void msr(ARMEmitter::SystemRegister reg, ARMEmitter::Register rt) {
|
||||
constexpr uint32_t Op = 0b1101'0101'0001 << 20;
|
||||
SystemRegisterMove(Op, rt, reg);
|
||||
}
|
||||
|
||||
void mrs(FEXCore::ARMEmitter::Register rd, FEXCore::ARMEmitter::SystemRegister reg) {
|
||||
void mrs(ARMEmitter::Register rd, ARMEmitter::SystemRegister reg) {
|
||||
constexpr uint32_t Op = 0b1101'0101'0011 << 20;
|
||||
SystemRegisterMove(Op, rd, reg);
|
||||
}
|
||||
@@ -130,7 +130,7 @@ private:
|
||||
}
|
||||
|
||||
// System instructions with register argument
|
||||
void SystemInstructionWithReg(uint32_t CRm, uint32_t op2, FEXCore::ARMEmitter::Register rt) {
|
||||
void SystemInstructionWithReg(uint32_t CRm, uint32_t op2, ARMEmitter::Register rt) {
|
||||
uint32_t Instr = 0b1101'0101'0000'0011'0001 << 12;
|
||||
|
||||
Instr |= CRm << 8;
|
||||
@@ -140,13 +140,13 @@ private:
|
||||
}
|
||||
|
||||
// Hints
|
||||
void Hint(FEXCore::ARMEmitter::HintRegister Reg) {
|
||||
void Hint(ARMEmitter::HintRegister Reg) {
|
||||
uint32_t Instr = 0b1101'0101'0000'0011'0010'0000'0001'1111U;
|
||||
Instr |= FEXCore::ToUnderlying(Reg);
|
||||
dc32(Instr);
|
||||
}
|
||||
// Barriers
|
||||
void Barrier(FEXCore::ARMEmitter::BarrierRegister Reg, uint32_t CRm) {
|
||||
void Barrier(ARMEmitter::BarrierRegister Reg, uint32_t CRm) {
|
||||
uint32_t Instr = 0b1101'0101'0000'0011'0011'0000'0001'1111U;
|
||||
Instr |= CRm << 8;
|
||||
Instr |= FEXCore::ToUnderlying(Reg);
|
||||
@@ -154,7 +154,7 @@ private:
|
||||
}
|
||||
|
||||
// System Instruction
|
||||
void SystemInstruction(uint32_t Op, uint32_t L, uint32_t SubOp, FEXCore::ARMEmitter::Register rt) {
|
||||
void SystemInstruction(uint32_t Op, uint32_t L, uint32_t SubOp, ARMEmitter::Register rt) {
|
||||
uint32_t Instr = Op;
|
||||
|
||||
Instr |= L << 21;
|
||||
@@ -165,7 +165,7 @@ private:
|
||||
}
|
||||
|
||||
// System register move
|
||||
void SystemRegisterMove(uint32_t Op, FEXCore::ARMEmitter::Register rt, FEXCore::ARMEmitter::SystemRegister reg) {
|
||||
void SystemRegisterMove(uint32_t Op, ARMEmitter::Register rt, ARMEmitter::SystemRegister reg) {
|
||||
uint32_t Instr = Op;
|
||||
|
||||
Instr |= FEXCore::ToUnderlying(reg);
|
||||
@@ -0,0 +1,311 @@
|
||||
// Collection of utilities from vixl.
|
||||
// Following is the vixl license.
|
||||
// Copyright 2015, VIXL authors
|
||||
// All rights reserved.
|
||||
//
|
||||
// Redistribution and use in source and binary forms, with or without
|
||||
// modification, are permitted provided that the following conditions are met:
|
||||
//
|
||||
// * Redistributions of source code must retain the above copyright notice,
|
||||
// this list of conditions and the following disclaimer.
|
||||
// * Redistributions in binary form must reproduce the above copyright notice,
|
||||
// this list of conditions and the following disclaimer in the documentation
|
||||
// and/or other materials provided with the distribution.
|
||||
// * Neither the name of ARM Limited nor the names of its contributors may be
|
||||
// used to endorse or promote products derived from this software without
|
||||
// specific prior written permission.
|
||||
//
|
||||
// THIS SOFTWARE IS PROVIDED BY THE COPYRIGHT HOLDERS CONTRIBUTORS "AS IS" AND
|
||||
// ANY EXPRESS OR IMPLIED WARRANTIES, INCLUDING, BUT NOT LIMITED TO, THE IMPLIED
|
||||
// WARRANTIES OF MERCHANTABILITY AND FITNESS FOR A PARTICULAR PURPOSE ARE
|
||||
// DISCLAIMED. IN NO EVENT SHALL THE COPYRIGHT OWNER OR CONTRIBUTORS BE LIABLE
|
||||
// FOR ANY DIRECT, INDIRECT, INCIDENTAL, SPECIAL, EXEMPLARY, OR CONSEQUENTIAL
|
||||
// DAMAGES (INCLUDING, BUT NOT LIMITED TO, PROCUREMENT OF SUBSTITUTE GOODS OR
|
||||
// SERVICES; LOSS OF USE, DATA, OR PROFITS; OR BUSINESS INTERRUPTION) HOWEVER
|
||||
// CAUSED AND ON ANY THEORY OF LIABILITY, WHETHER IN CONTRACT, STRICT LIABILITY,
|
||||
// OR TORT (INCLUDING NEGLIGENCE OR OTHERWISE) ARISING IN ANY WAY OUT OF THE USE
|
||||
// OF THIS SOFTWARE, EVEN IF ADVISED OF THE POSSIBILITY OF SUCH DAMAGE.
|
||||
|
||||
|
||||
// Test if a given value can be encoded in the immediate field of a logical
|
||||
// instruction.
|
||||
// If it can be encoded, the function returns true, and values pointed to by n,
|
||||
// imm_s and imm_r are updated with immediates encoded in the format required
|
||||
// by the corresponding fields in the logical instruction.
|
||||
// If it can not be encoded, the function returns false, and the values pointed
|
||||
// to by n, imm_s and imm_r are undefined.
|
||||
static bool IsImmLogical(uint64_t value,
|
||||
unsigned width,
|
||||
unsigned* n,
|
||||
unsigned* imm_s,
|
||||
unsigned* imm_r) {
|
||||
[[maybe_unused]] constexpr auto kBRegSize = 8;
|
||||
[[maybe_unused]] constexpr auto kHRegSize = 16;
|
||||
[[maybe_unused]] constexpr auto kSRegSize = 32;
|
||||
[[maybe_unused]] constexpr auto kDRegSize = 64;
|
||||
|
||||
constexpr auto kWRegSize = 32;
|
||||
constexpr auto kXRegSize = 64;
|
||||
|
||||
LOGMAN_THROW_A_FMT((width == kBRegSize) || (width == kHRegSize) ||
|
||||
(width == kSRegSize) || (width == kDRegSize), "Unexpected imm size");
|
||||
|
||||
bool negate = false;
|
||||
|
||||
// Logical immediates are encoded using parameters n, imm_s and imm_r using
|
||||
// the following table:
|
||||
//
|
||||
// N imms immr size S R
|
||||
// 1 ssssss rrrrrr 64 UInt(ssssss) UInt(rrrrrr)
|
||||
// 0 0sssss xrrrrr 32 UInt(sssss) UInt(rrrrr)
|
||||
// 0 10ssss xxrrrr 16 UInt(ssss) UInt(rrrr)
|
||||
// 0 110sss xxxrrr 8 UInt(sss) UInt(rrr)
|
||||
// 0 1110ss xxxxrr 4 UInt(ss) UInt(rr)
|
||||
// 0 11110s xxxxxr 2 UInt(s) UInt(r)
|
||||
// (s bits must not be all set)
|
||||
//
|
||||
// A pattern is constructed of size bits, where the least significant S+1 bits
|
||||
// are set. The pattern is rotated right by R, and repeated across a 32 or
|
||||
// 64-bit value, depending on destination register width.
|
||||
//
|
||||
// Put another way: the basic format of a logical immediate is a single
|
||||
// contiguous stretch of 1 bits, repeated across the whole word at intervals
|
||||
// given by a power of 2. To identify them quickly, we first locate the
|
||||
// lowest stretch of 1 bits, then the next 1 bit above that; that combination
|
||||
// is different for every logical immediate, so it gives us all the
|
||||
// information we need to identify the only logical immediate that our input
|
||||
// could be, and then we simply check if that's the value we actually have.
|
||||
//
|
||||
// (The rotation parameter does give the possibility of the stretch of 1 bits
|
||||
// going 'round the end' of the word. To deal with that, we observe that in
|
||||
// any situation where that happens the bitwise NOT of the value is also a
|
||||
// valid logical immediate. So we simply invert the input whenever its low bit
|
||||
// is set, and then we know that the rotated case can't arise.)
|
||||
|
||||
if (value & 1) {
|
||||
// If the low bit is 1, negate the value, and set a flag to remember that we
|
||||
// did (so that we can adjust the return values appropriately).
|
||||
negate = true;
|
||||
value = ~value;
|
||||
}
|
||||
|
||||
if (width <= kWRegSize) {
|
||||
// To handle 8/16/32-bit logical immediates, the very easiest thing is to repeat
|
||||
// the input value to fill a 64-bit word. The correct encoding of that as a
|
||||
// logical immediate will also be the correct encoding of the value.
|
||||
|
||||
// Avoid making the assumption that the most-significant 56/48/32 bits are zero by
|
||||
// shifting the value left and duplicating it.
|
||||
for (unsigned bits = width; bits <= kWRegSize; bits *= 2) {
|
||||
value <<= bits;
|
||||
uint64_t mask = (UINT64_C(1) << bits) - 1;
|
||||
value |= ((value >> bits) & mask);
|
||||
}
|
||||
}
|
||||
|
||||
// The basic analysis idea: imagine our input word looks like this.
|
||||
//
|
||||
// 0011111000111110001111100011111000111110001111100011111000111110
|
||||
// c b a
|
||||
// |<--d-->|
|
||||
//
|
||||
// We find the lowest set bit (as an actual power-of-2 value, not its index)
|
||||
// and call it a. Then we add a to our original number, which wipes out the
|
||||
// bottommost stretch of set bits and replaces it with a 1 carried into the
|
||||
// next zero bit. Then we look for the new lowest set bit, which is in
|
||||
// position b, and subtract it, so now our number is just like the original
|
||||
// but with the lowest stretch of set bits completely gone. Now we find the
|
||||
// lowest set bit again, which is position c in the diagram above. Then we'll
|
||||
// measure the distance d between bit positions a and c (using CLZ), and that
|
||||
// tells us that the only valid logical immediate that could possibly be equal
|
||||
// to this number is the one in which a stretch of bits running from a to just
|
||||
// below b is replicated every d bits.
|
||||
uint64_t a = LowestSetBit(value);
|
||||
uint64_t value_plus_a = value + a;
|
||||
uint64_t b = LowestSetBit(value_plus_a);
|
||||
uint64_t value_plus_a_minus_b = value_plus_a - b;
|
||||
uint64_t c = LowestSetBit(value_plus_a_minus_b);
|
||||
|
||||
int d, clz_a, out_n;
|
||||
uint64_t mask;
|
||||
|
||||
if (c != 0) {
|
||||
// The general case, in which there is more than one stretch of set bits.
|
||||
// Compute the repeat distance d, and set up a bitmask covering the basic
|
||||
// unit of repetition (i.e. a word with the bottom d bits set). Also, in all
|
||||
// of these cases the N bit of the output will be zero.
|
||||
clz_a = CountLeadingZeros(a, kXRegSize);
|
||||
int clz_c = CountLeadingZeros(c, kXRegSize);
|
||||
d = clz_a - clz_c;
|
||||
mask = ((UINT64_C(1) << d) - 1);
|
||||
out_n = 0;
|
||||
} else {
|
||||
// Handle degenerate cases.
|
||||
//
|
||||
// If any of those 'find lowest set bit' operations didn't find a set bit at
|
||||
// all, then the word will have been zero thereafter, so in particular the
|
||||
// last lowest_set_bit operation will have returned zero. So we can test for
|
||||
// all the special case conditions in one go by seeing if c is zero.
|
||||
if (a == 0) {
|
||||
// The input was zero (or all 1 bits, which will come to here too after we
|
||||
// inverted it at the start of the function), for which we just return
|
||||
// false.
|
||||
return false;
|
||||
} else {
|
||||
// Otherwise, if c was zero but a was not, then there's just one stretch
|
||||
// of set bits in our word, meaning that we have the trivial case of
|
||||
// d == 64 and only one 'repetition'. Set up all the same variables as in
|
||||
// the general case above, and set the N bit in the output.
|
||||
clz_a = CountLeadingZeros(a, kXRegSize);
|
||||
d = 64;
|
||||
mask = ~UINT64_C(0);
|
||||
out_n = 1;
|
||||
}
|
||||
}
|
||||
|
||||
// If the repeat period d is not a power of two, it can't be encoded.
|
||||
if (!IsPowerOf2(d)) {
|
||||
return false;
|
||||
}
|
||||
|
||||
if (((b - a) & ~mask) != 0) {
|
||||
// If the bit stretch (b - a) does not fit within the mask derived from the
|
||||
// repeat period, then fail.
|
||||
return false;
|
||||
}
|
||||
|
||||
// The only possible option is b - a repeated every d bits. Now we're going to
|
||||
// actually construct the valid logical immediate derived from that
|
||||
// specification, and see if it equals our original input.
|
||||
//
|
||||
// To repeat a value every d bits, we multiply it by a number of the form
|
||||
// (1 + 2^d + 2^(2d) + ...), i.e. 0x0001000100010001 or similar. These can
|
||||
// be derived using a table lookup on CLZ(d).
|
||||
static const uint64_t multipliers[] = {
|
||||
0x0000000000000001UL,
|
||||
0x0000000100000001UL,
|
||||
0x0001000100010001UL,
|
||||
0x0101010101010101UL,
|
||||
0x1111111111111111UL,
|
||||
0x5555555555555555UL,
|
||||
};
|
||||
uint64_t multiplier = multipliers[CountLeadingZeros(d, kXRegSize) - 57];
|
||||
uint64_t candidate = (b - a) * multiplier;
|
||||
|
||||
if (value != candidate) {
|
||||
// The candidate pattern doesn't match our input value, so fail.
|
||||
return false;
|
||||
}
|
||||
|
||||
// We have a match! This is a valid logical immediate, so now we have to
|
||||
// construct the bits and pieces of the instruction encoding that generates
|
||||
// it.
|
||||
|
||||
// Count the set bits in our basic stretch. The special case of clz(0) == -1
|
||||
// makes the answer come out right for stretches that reach the very top of
|
||||
// the word (e.g. numbers like 0xffffc00000000000).
|
||||
int clz_b = (b == 0) ? -1 : CountLeadingZeros(b, kXRegSize);
|
||||
int s = clz_a - clz_b;
|
||||
|
||||
// Decide how many bits to rotate right by, to put the low bit of that basic
|
||||
// stretch in position a.
|
||||
int r;
|
||||
if (negate) {
|
||||
// If we inverted the input right at the start of this function, here's
|
||||
// where we compensate: the number of set bits becomes the number of clear
|
||||
// bits, and the rotation count is based on position b rather than position
|
||||
// a (since b is the location of the 'lowest' 1 bit after inversion).
|
||||
s = d - s;
|
||||
r = (clz_b + 1) & (d - 1);
|
||||
} else {
|
||||
r = (clz_a + 1) & (d - 1);
|
||||
}
|
||||
|
||||
// Now we're done, except for having to encode the S output in such a way that
|
||||
// it gives both the number of set bits and the length of the repeated
|
||||
// segment. The s field is encoded like this:
|
||||
//
|
||||
// imms size S
|
||||
// ssssss 64 UInt(ssssss)
|
||||
// 0sssss 32 UInt(sssss)
|
||||
// 10ssss 16 UInt(ssss)
|
||||
// 110sss 8 UInt(sss)
|
||||
// 1110ss 4 UInt(ss)
|
||||
// 11110s 2 UInt(s)
|
||||
//
|
||||
// So we 'or' (2 * -d) with our computed s to form imms.
|
||||
if ((n != NULL) || (imm_s != NULL) || (imm_r != NULL)) {
|
||||
*n = out_n;
|
||||
*imm_s = ((2 * -d) | (s - 1)) & 0x3f;
|
||||
*imm_r = r;
|
||||
}
|
||||
|
||||
return true;
|
||||
}
|
||||
|
||||
private:
|
||||
|
||||
template <typename V>
|
||||
static inline bool IsPowerOf2(V value) {
|
||||
return (value != 0) && ((value & (value - 1)) == 0);
|
||||
}
|
||||
|
||||
// Some compilers dislike negating unsigned integers,
|
||||
// so we provide an equivalent.
|
||||
template <typename T>
|
||||
static inline T UnsignedNegate(T value) {
|
||||
static_assert(std::is_unsigned<T>::value);
|
||||
return ~value + 1;
|
||||
}
|
||||
|
||||
static inline uint64_t LowestSetBit(uint64_t value) {
|
||||
return value & UnsignedNegate(value);
|
||||
}
|
||||
|
||||
template <typename V>
|
||||
static inline int CountLeadingZeros(V value, int width = (sizeof(V) * 8)) {
|
||||
#if COMPILER_HAS_BUILTIN_CLZ
|
||||
if (width == 32) {
|
||||
return (value == 0) ? 32 : __builtin_clz(static_cast<unsigned>(value));
|
||||
} else if (width == 64) {
|
||||
return (value == 0) ? 64 : __builtin_clzll(value);
|
||||
}
|
||||
#endif
|
||||
return CountLeadingZerosFallBack(value, width);
|
||||
}
|
||||
|
||||
static inline int CountLeadingZerosFallBack(uint64_t value, int width) {
|
||||
LOGMAN_THROW_A_FMT(IsPowerOf2(width) && (width <= 64), "Invalid width");
|
||||
if (value == 0) {
|
||||
return width;
|
||||
}
|
||||
int count = 0;
|
||||
value = value << (64 - width);
|
||||
if ((value & UINT64_C(0xffffffff00000000)) == 0) {
|
||||
count += 32;
|
||||
value = value << 32;
|
||||
}
|
||||
if ((value & UINT64_C(0xffff000000000000)) == 0) {
|
||||
count += 16;
|
||||
value = value << 16;
|
||||
}
|
||||
if ((value & UINT64_C(0xff00000000000000)) == 0) {
|
||||
count += 8;
|
||||
value = value << 8;
|
||||
}
|
||||
if ((value & UINT64_C(0xf000000000000000)) == 0) {
|
||||
count += 4;
|
||||
value = value << 4;
|
||||
}
|
||||
if ((value & UINT64_C(0xc000000000000000)) == 0) {
|
||||
count += 2;
|
||||
value = value << 2;
|
||||
}
|
||||
if ((value & UINT64_C(0x8000000000000000)) == 0) {
|
||||
count += 1;
|
||||
}
|
||||
count += (value == 0);
|
||||
return count;
|
||||
}
|
||||
|
||||
public:
|
||||
+15
-73
@@ -20,8 +20,7 @@ the coding style of LLVM. It can also be installed as a pre-commit git hook to
|
||||
check the coding style before submitting it. The canonical source of this script
|
||||
is in the LLVM source tree under llvm/utils/git.
|
||||
|
||||
For C/C++ code it uses clang-format and for Python code it uses darker (which
|
||||
in turn invokes black).
|
||||
For C/C++ code it uses clang-format.
|
||||
|
||||
You can learn more about the LLVM coding style on llvm.org:
|
||||
https://llvm.org/docs/CodingStandards.html
|
||||
@@ -31,8 +30,8 @@ directory:
|
||||
|
||||
ln -s $(pwd)/llvm/utils/git/code-format-helper.py .git/hooks/pre-commit
|
||||
|
||||
You can control the exact path to clang-format or darker with the following
|
||||
environment variables: $CLANG_FORMAT_PATH and $DARKER_FORMAT_PATH.
|
||||
You can control the exact path to clang-format with the following
|
||||
environment variable: $CLANG_FORMAT_PATH.
|
||||
"""
|
||||
|
||||
|
||||
@@ -173,6 +172,12 @@ class ClangFormatHelper(FormatHelper):
|
||||
name = "clang-format"
|
||||
friendly_name = "C/C++ code formatter"
|
||||
|
||||
@property
|
||||
def cformat_wrapper_path(self) -> str:
|
||||
relpath = "../../Scripts/clang-format.py"
|
||||
curpath = os.path.dirname(os.path.abspath(__file__))
|
||||
return os.path.abspath(os.path.normpath(os.path.join(curpath, relpath)))
|
||||
|
||||
@property
|
||||
def instructions(self) -> str:
|
||||
return " ".join(self.cf_cmd)
|
||||
@@ -210,7 +215,11 @@ class ClangFormatHelper(FormatHelper):
|
||||
if not cpp_files:
|
||||
return None
|
||||
|
||||
cf_cmd = [self.clang_fmt_path, "--diff"]
|
||||
cf_cmd = [
|
||||
self.clang_fmt_path,
|
||||
f"--binary={self.cformat_wrapper_path}",
|
||||
"--diff",
|
||||
]
|
||||
|
||||
if args.start_rev and args.end_rev:
|
||||
cf_cmd.append(args.start_rev)
|
||||
@@ -235,74 +244,7 @@ class ClangFormatHelper(FormatHelper):
|
||||
else:
|
||||
return None
|
||||
|
||||
|
||||
class DarkerFormatHelper(FormatHelper):
|
||||
name = "darker"
|
||||
friendly_name = "Python code formatter"
|
||||
|
||||
@property
|
||||
def instructions(self) -> str:
|
||||
return " ".join(self.darker_cmd)
|
||||
|
||||
def filter_changed_files(self, changed_files: List[str]) -> List[str]:
|
||||
filtered_files = []
|
||||
for path in changed_files:
|
||||
name, ext = os.path.splitext(path)
|
||||
if ext == ".py":
|
||||
filtered_files.append(path)
|
||||
|
||||
return filtered_files
|
||||
|
||||
@property
|
||||
def darker_fmt_path(self) -> str:
|
||||
if "DARKER_FORMAT_PATH" in os.environ:
|
||||
return os.environ["DARKER_FORMAT_PATH"]
|
||||
return "darker"
|
||||
|
||||
def has_tool(self) -> bool:
|
||||
cmd = [self.darker_fmt_path, "--version"]
|
||||
proc = None
|
||||
try:
|
||||
proc = subprocess.run(cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE)
|
||||
except:
|
||||
return False
|
||||
return proc.returncode == 0
|
||||
|
||||
def format_run(self, changed_files: List[str], args: FormatArgs) -> Optional[str]:
|
||||
py_files = self.filter_changed_files(changed_files)
|
||||
if not py_files:
|
||||
return None
|
||||
darker_cmd = [
|
||||
self.darker_fmt_path,
|
||||
"--check",
|
||||
"--diff",
|
||||
]
|
||||
if args.start_rev and args.end_rev:
|
||||
darker_cmd += ["-r", f"{args.start_rev}...{args.end_rev}"]
|
||||
darker_cmd += py_files
|
||||
if args.verbose:
|
||||
print(f"Running: {' '.join(darker_cmd)}")
|
||||
self.darker_cmd = darker_cmd
|
||||
proc = subprocess.run(
|
||||
darker_cmd, stdout=subprocess.PIPE, stderr=subprocess.PIPE
|
||||
)
|
||||
if args.verbose:
|
||||
sys.stdout.write(proc.stderr.decode("utf-8"))
|
||||
|
||||
if proc.returncode != 0:
|
||||
# formatting needed, or the command otherwise failed
|
||||
if args.verbose:
|
||||
print(f"error: {self.name} exited with code {proc.returncode}")
|
||||
# Print the diff in the log so that it is viewable there
|
||||
print(proc.stdout.decode("utf-8"))
|
||||
return proc.stdout.decode("utf-8")
|
||||
else:
|
||||
sys.stdout.write(proc.stdout.decode("utf-8"))
|
||||
return None
|
||||
|
||||
|
||||
ALL_FORMATTERS = (DarkerFormatHelper(), ClangFormatHelper())
|
||||
|
||||
ALL_FORMATTERS = [ClangFormatHelper()]
|
||||
|
||||
def hook_main():
|
||||
# fill out args
|
||||
|
||||
Vendored
+1
-1
Submodule External/vixl updated: 7725aec177...a90f5d5020.
@@ -55,6 +55,7 @@ class OpDefinition:
|
||||
DynamicDispatch: bool
|
||||
JITDispatch: bool
|
||||
JITDispatchOverride: str
|
||||
TiedSource: int
|
||||
Arguments: list
|
||||
EmitValidation: list
|
||||
Desc: list
|
||||
@@ -77,6 +78,7 @@ class OpDefinition:
|
||||
self.DynamicDispatch = False
|
||||
self.JITDispatch = True
|
||||
self.JITDispatchOverride = None
|
||||
self.TiedSource = -1
|
||||
self.Arguments = []
|
||||
self.EmitValidation = []
|
||||
self.Desc = []
|
||||
@@ -248,6 +250,9 @@ def parse_ops(ops):
|
||||
if "JITDispatchOverride" in op_val:
|
||||
OpDef.JITDispatchOverride = op_val["JITDispatchOverride"]
|
||||
|
||||
if "TiedSource" in op_val:
|
||||
OpDef.TiedSource = op_val["TiedSource"]
|
||||
|
||||
# Do some fixups of the data here
|
||||
if len(OpDef.EmitValidation) != 0:
|
||||
for i in range(len(OpDef.EmitValidation)):
|
||||
@@ -372,13 +377,30 @@ def print_ir_sizes():
|
||||
|
||||
output_file.write("[[maybe_unused, nodiscard]] static size_t GetSize(IROps Op) { return IRSizes[Op]; }\n\n")
|
||||
|
||||
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] std::string_view const& GetName(IROps Op);\n")
|
||||
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] uint8_t GetArgs(IROps Op);\n")
|
||||
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] uint8_t GetRAArgs(IROps Op);\n")
|
||||
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] FEXCore::IR::RegisterClassType GetRegClass(IROps Op);\n\n")
|
||||
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] bool HasSideEffects(IROps Op);\n")
|
||||
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] bool ImplicitFlagClobber(IROps Op);\n")
|
||||
output_file.write("[[nodiscard, gnu::const, gnu::visibility(\"default\")]] bool GetHasDest(IROps Op);\n")
|
||||
output_file.write(
|
||||
'[[nodiscard, gnu::const, gnu::visibility("default")]] std::string_view const& GetName(IROps Op);\n'
|
||||
)
|
||||
output_file.write(
|
||||
'[[nodiscard, gnu::const, gnu::visibility("default")]] uint8_t GetArgs(IROps Op);\n'
|
||||
)
|
||||
output_file.write(
|
||||
'[[nodiscard, gnu::const, gnu::visibility("default")]] uint8_t GetRAArgs(IROps Op);\n'
|
||||
)
|
||||
output_file.write(
|
||||
'[[nodiscard, gnu::const, gnu::visibility("default")]] FEXCore::IR::RegisterClassType GetRegClass(IROps Op);\n\n'
|
||||
)
|
||||
output_file.write(
|
||||
'[[nodiscard, gnu::const, gnu::visibility("default")]] bool HasSideEffects(IROps Op);\n'
|
||||
)
|
||||
output_file.write(
|
||||
'[[nodiscard, gnu::const, gnu::visibility("default")]] bool ImplicitFlagClobber(IROps Op);\n'
|
||||
)
|
||||
output_file.write(
|
||||
'[[nodiscard, gnu::const, gnu::visibility("default")]] bool GetHasDest(IROps Op);\n'
|
||||
)
|
||||
output_file.write(
|
||||
'[[nodiscard, gnu::const, gnu::visibility("default")]] int8_t TiedSource(IROps Op);\n'
|
||||
)
|
||||
|
||||
output_file.write("#undef IROP_SIZES\n")
|
||||
output_file.write("#endif\n\n")
|
||||
@@ -471,15 +493,25 @@ def print_ir_getraargs():
|
||||
def print_ir_hassideeffects():
|
||||
output_file.write("#ifdef IROP_HASSIDEEFFECTS_IMPL\n")
|
||||
|
||||
for array, prop in [("SideEffects", "HasSideEffects"),
|
||||
("ImplicitFlagClobbers", "ImplicitFlagClobber")]:
|
||||
output_file.write(f"constexpr std::array<uint8_t, OP_LAST + 1> {array} = {{\n")
|
||||
for array, prop, T in [
|
||||
("SideEffects", "HasSideEffects", "bool"),
|
||||
("ImplicitFlagClobbers", "ImplicitFlagClobber", "bool"),
|
||||
("TiedSources", "TiedSource", "int8_t"),
|
||||
]:
|
||||
output_file.write(
|
||||
f"constexpr std::array<{'uint8_t' if T == 'bool' else T}, OP_LAST + 1> {array} = {{\n"
|
||||
)
|
||||
for op in IROps:
|
||||
output_file.write("\t{},\n".format(("true" if getattr(op, prop) else "false")))
|
||||
if T == "bool":
|
||||
output_file.write(
|
||||
"\t{},\n".format(("true" if getattr(op, prop) else "false"))
|
||||
)
|
||||
else:
|
||||
output_file.write(f"\t{getattr(op, prop)},\n")
|
||||
|
||||
output_file.write("};\n\n")
|
||||
|
||||
output_file.write(f"bool {prop}(IROps Op) {{\n")
|
||||
output_file.write(f"{T} {prop}(IROps Op) {{\n")
|
||||
output_file.write(f" return {array}[Op];\n")
|
||||
output_file.write("}\n")
|
||||
|
||||
@@ -520,14 +552,20 @@ def print_ir_arg_printer():
|
||||
output_file.write("\t*out << \" \";\n")
|
||||
|
||||
SSAArgNum = 0
|
||||
FirstArg = True
|
||||
for i in range(0, len(op.Arguments)):
|
||||
arg = op.Arguments[i]
|
||||
LastArg = len(op.Arguments) - i - 1 == 0
|
||||
|
||||
# No point printing temporaries that we can't recover
|
||||
if arg.Temporary:
|
||||
# Temporary that we can't recover
|
||||
output_file.write("\t*out << \"{}:Tmp:{}\";\n".format(arg.Type, arg.Name))
|
||||
elif arg.IsSSA:
|
||||
continue
|
||||
|
||||
if FirstArg:
|
||||
FirstArg = False
|
||||
else:
|
||||
output_file.write('\t*out << ", ";\n')
|
||||
|
||||
if arg.IsSSA:
|
||||
# SSA value
|
||||
output_file.write("\tPrintArg(out, IR, Op->Header.Args[{}], RAData);\n".format(SSAArgNum))
|
||||
SSAArgNum = SSAArgNum + 1
|
||||
@@ -535,9 +573,6 @@ def print_ir_arg_printer():
|
||||
# User defined op that is stored
|
||||
output_file.write("\tPrintArg(out, IR, Op->{});\n".format(arg.Name))
|
||||
|
||||
if not LastArg:
|
||||
output_file.write("\t*out << \", \";\n")
|
||||
|
||||
output_file.write("break;\n")
|
||||
output_file.write("}\n")
|
||||
|
||||
|
||||
@@ -95,6 +95,7 @@ set (SRCS
|
||||
Interface/Core/ObjectCache/JobHandling.cpp
|
||||
Interface/Core/ObjectCache/NamedRegionObjectHandler.cpp
|
||||
Interface/Core/ObjectCache/ObjectCacheService.cpp
|
||||
Interface/Core/OpcodeDispatcher/AVX_128.cpp
|
||||
Interface/Core/OpcodeDispatcher/Crypto.cpp
|
||||
Interface/Core/OpcodeDispatcher/Flags.cpp
|
||||
Interface/Core/OpcodeDispatcher/Vector.cpp
|
||||
@@ -134,22 +135,16 @@ set (SRCS
|
||||
Interface/GDBJIT/GDBJIT.cpp
|
||||
Interface/IR/AOTIR.cpp
|
||||
Interface/IR/IRDumper.cpp
|
||||
Interface/IR/IRParser.cpp
|
||||
Interface/IR/IREmitter.cpp
|
||||
Interface/IR/PassManager.cpp
|
||||
Interface/IR/Passes/ConstProp.cpp
|
||||
Interface/IR/Passes/DeadCodeElimination.cpp
|
||||
Interface/IR/Passes/DeadContextStoreElimination.cpp
|
||||
Interface/IR/Passes/IRCompaction.cpp
|
||||
Interface/IR/Passes/IRDumperPass.cpp
|
||||
Interface/IR/Passes/IRValidation.cpp
|
||||
Interface/IR/Passes/RAValidation.cpp
|
||||
Interface/IR/Passes/LongDivideRemovalPass.cpp
|
||||
Interface/IR/Passes/ValueDominanceValidation.cpp
|
||||
Interface/IR/Passes/RedundantFlagCalculationElimination.cpp
|
||||
Interface/IR/Passes/DeadStoreElimination.cpp
|
||||
Interface/IR/Passes/RegisterAllocationPass.cpp
|
||||
Interface/IR/Passes/InlineCallOptimization.cpp
|
||||
Utils/Telemetry.cpp
|
||||
Utils/Threads.cpp
|
||||
Utils/Profiler.cpp
|
||||
@@ -194,7 +189,7 @@ endif()
|
||||
# Some defines for the softfloat library
|
||||
list(APPEND DEFINES "-DSOFTFLOAT_BUILTIN_CLZ")
|
||||
|
||||
set (LIBS fmt::fmt vixl xxHash::xxhash FEXHeaderUtils)
|
||||
set (LIBS fmt::fmt vixl xxHash::xxhash FEXHeaderUtils CodeEmitter)
|
||||
|
||||
if (NOT MINGW_BUILD)
|
||||
list (APPEND LIBS dl)
|
||||
@@ -373,16 +368,6 @@ function(AddLibrary Name Type)
|
||||
target_link_libraries(${Name} FEXCore_Base)
|
||||
target_compile_options(${Name} PRIVATE ${FEX_TUNE_COMPILE_FLAGS})
|
||||
set_target_properties(${Name} PROPERTIES OUTPUT_NAME FEXCore)
|
||||
if (MINGW_BUILD)
|
||||
# Mingw build isn't building a linux shared library, so it can't have a SONAME.
|
||||
set_target_properties(${Name} PROPERTIES NO_SONAME ON)
|
||||
# Change the suffixes otherwise cmake continues using .a and .so
|
||||
if (${Type} STREQUAL SHARED)
|
||||
set_target_properties(${Name} PROPERTIES SUFFIX ".dll")
|
||||
elseif(${Type} STREQUAL STATIC)
|
||||
set_target_properties(${Name} PROPERTIES SUFFIX ".lib")
|
||||
endif()
|
||||
endif()
|
||||
|
||||
AddDefaultOptionsToTarget(${Name})
|
||||
endfunction()
|
||||
|
||||
@@ -20,12 +20,12 @@ struct BitSet final {
|
||||
|
||||
ElementType* Memory;
|
||||
void Allocate(size_t Elements) {
|
||||
size_t AllocateSize = AlignUp(Elements, MinimumSizeBits) / MinimumSize;
|
||||
size_t AllocateSize = ToBytes(Elements);
|
||||
LOGMAN_THROW_AA_FMT((AllocateSize * MinimumSize) >= Elements, "Fail");
|
||||
Memory = static_cast<ElementType*>(FEXCore::Allocator::malloc(AllocateSize));
|
||||
}
|
||||
void Realloc(size_t Elements) {
|
||||
size_t AllocateSize = AlignUp(Elements, MinimumSizeBits) / MinimumSize;
|
||||
size_t AllocateSize = ToBytes(Elements);
|
||||
LOGMAN_THROW_AA_FMT((AllocateSize * MinimumSize) >= Elements, "Fail");
|
||||
Memory = static_cast<ElementType*>(FEXCore::Allocator::realloc(Memory, AllocateSize));
|
||||
}
|
||||
@@ -43,10 +43,13 @@ struct BitSet final {
|
||||
Memory[Element / MinimumSizeBits] &= (1ULL << (Element % MinimumSizeBits));
|
||||
}
|
||||
void MemClear(size_t Elements) {
|
||||
memset(Memory, 0, AlignUp(Elements / MinimumSizeBits, MinimumSizeBits));
|
||||
memset(Memory, 0, ToBytes(Elements));
|
||||
}
|
||||
void MemSet(size_t Elements) {
|
||||
memset(Memory, 0xFF, AlignUp(Elements / MinimumSizeBits, MinimumSizeBits));
|
||||
memset(Memory, 0xFF, ToBytes(Elements));
|
||||
}
|
||||
uint32_t ToBytes(size_t Elements) {
|
||||
return AlignUp(Elements, MinimumSizeBits) / MinimumSize;
|
||||
}
|
||||
|
||||
// This very explicitly doesn't let you take an address
|
||||
|
||||
@@ -50,8 +50,6 @@
|
||||
"DISABLESVE": "disablesve",
|
||||
"ENABLEAVX": "enableavx",
|
||||
"DISABLEAVX": "disableavx",
|
||||
"ENABLEAVX2": "enableavx2",
|
||||
"DISABLEAVX2": "disableavx2",
|
||||
"ENABLEAFP": "enableafp",
|
||||
"DISABLEAFP": "disableafp",
|
||||
"ENABLELRCPC": "enablelrcpc",
|
||||
@@ -86,7 +84,6 @@
|
||||
"\toff: Default CPU features queried from CPU features",
|
||||
"\t{enable,disable}sve: Will force enable or disable sve even if the host doesn't support it",
|
||||
"\t{enable,disable}avx: Will force enable or disable avx even if the host doesn't support it",
|
||||
"\t{enable,disable}avx2: Will force enable or disable avx2 even if the host doesn't support it",
|
||||
"\t{enable,disable}afp: Will force enable or disable afp even if the host doesn't support it",
|
||||
"\t{enable,disable}lrcpc: Will force enable or disable lrcpc even if the host doesn't support it",
|
||||
"\t{enable,disable}lrcpc2: Will force enable or disable lrcpc2 even if the host doesn't support it",
|
||||
@@ -401,14 +398,14 @@
|
||||
},
|
||||
"VectorTSOEnabled": {
|
||||
"Type": "bool",
|
||||
"Default": "true",
|
||||
"Default": "false",
|
||||
"Desc": [
|
||||
"When TSO emulation is enabled, controls if vector loadstores should also be atomic."
|
||||
]
|
||||
},
|
||||
"MemcpySetTSOEnabled": {
|
||||
"Type": "bool",
|
||||
"Default": "true",
|
||||
"Default": "false",
|
||||
"Desc": [
|
||||
"When TSO emulation is enabled, controls if memcpy and memset should also be atomic.",
|
||||
"Only affects REP MOVS and REP STOS instructions"
|
||||
|
||||
@@ -57,6 +57,7 @@ namespace HLE {
|
||||
|
||||
namespace FEXCore::IR {
|
||||
class RegisterAllocationData;
|
||||
struct IRListCopy;
|
||||
class IRListView;
|
||||
namespace Validation {
|
||||
class IRValidation;
|
||||
@@ -103,6 +104,9 @@ public:
|
||||
uint32_t ReconstructCompactedEFLAGS(FEXCore::Core::InternalThreadState* Thread, bool WasInJIT, uint64_t* HostGPRs, uint64_t PSTATE) override;
|
||||
void SetFlagsFromCompactedEFLAGS(FEXCore::Core::InternalThreadState* Thread, uint32_t EFLAGS) override;
|
||||
|
||||
void ReconstructXMMRegisters(const FEXCore::Core::InternalThreadState* Thread, __uint128_t* XMM_Low, __uint128_t* YMM_High) override;
|
||||
void SetXMMRegistersFromState(FEXCore::Core::InternalThreadState* Thread, const __uint128_t* XMM_Low, const __uint128_t* YMM_High) override;
|
||||
|
||||
/**
|
||||
* @brief Used to create FEX thread objects in preparation for creating a true OS thread. Does set a TID or PID.
|
||||
*
|
||||
@@ -205,9 +209,6 @@ public:
|
||||
CoreRunningMode RunningMode {CoreRunningMode::MODE_RUN};
|
||||
uint64_t VirtualMemSize {1ULL << 36};
|
||||
|
||||
// this is for internal use
|
||||
bool ValidateIRarser {false};
|
||||
|
||||
// Used if the JIT needs to have its interrupt fault code emitted.
|
||||
bool NeedsPendingInterruptFaultCheck {false};
|
||||
|
||||
@@ -293,8 +294,7 @@ public:
|
||||
void RemoveCustomIREntrypoint(uintptr_t Entrypoint);
|
||||
|
||||
struct GenerateIRResult {
|
||||
FEXCore::IR::IRListView* IRList;
|
||||
FEXCore::IR::RegisterAllocationData::UniquePtr RAData;
|
||||
fextl::unique_ptr<FEXCore::IR::IRStorageBase> IR;
|
||||
uint64_t TotalInstructions;
|
||||
uint64_t TotalInstructionsLength;
|
||||
uint64_t StartAddr;
|
||||
@@ -305,9 +305,8 @@ public:
|
||||
|
||||
struct CompileCodeResult {
|
||||
void* CompiledCode;
|
||||
FEXCore::IR::IRListView* IRData;
|
||||
fextl::unique_ptr<FEXCore::IR::IRStorageBase> IR;
|
||||
FEXCore::Core::DebugData* DebugData;
|
||||
FEXCore::IR::RegisterAllocationData::UniquePtr RAData;
|
||||
bool GeneratedIR;
|
||||
uint64_t StartAddr;
|
||||
uint64_t Length;
|
||||
|
||||
@@ -2,8 +2,6 @@
|
||||
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
|
||||
#include "FEXCore/Core/X86Enums.h"
|
||||
#include "FEXCore/Utils/AllocatorHooks.h"
|
||||
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
|
||||
#include "Interface/Core/ArchHelpers/CodeEmitter/Registers.h"
|
||||
#include "Interface/Core/Dispatcher/Dispatcher.h"
|
||||
#include "Interface/Context/Context.h"
|
||||
#include "Interface/HLE/Thunks/Thunks.h"
|
||||
@@ -13,6 +11,8 @@
|
||||
#include <FEXCore/Utils/MathUtils.h>
|
||||
|
||||
#include <FEXHeaderUtils/BitUtils.h>
|
||||
#include <CodeEmitter/Emitter.h>
|
||||
#include <CodeEmitter/Registers.h>
|
||||
|
||||
#include <aarch64/cpu-aarch64.h>
|
||||
#include <aarch64/instructions-aarch64.h>
|
||||
@@ -24,117 +24,103 @@
|
||||
#include <utility>
|
||||
|
||||
namespace FEXCore::CPU {
|
||||
// Register x18 is unused in the current configuration.
|
||||
// This is due to it being a platform register on wine platforms.
|
||||
// TODO: Allow x18 register allocation on Linux in the future to gain one more register.
|
||||
|
||||
namespace x64 {
|
||||
#ifndef _M_ARM_64EC
|
||||
// All but x19 and x29 are caller saved
|
||||
constexpr std::array<FEXCore::ARMEmitter::Register, 18> SRA = {
|
||||
FEXCore::ARMEmitter::Reg::r4,
|
||||
FEXCore::ARMEmitter::Reg::r5,
|
||||
FEXCore::ARMEmitter::Reg::r6,
|
||||
FEXCore::ARMEmitter::Reg::r7,
|
||||
FEXCore::ARMEmitter::Reg::r8,
|
||||
FEXCore::ARMEmitter::Reg::r9,
|
||||
FEXCore::ARMEmitter::Reg::r10,
|
||||
FEXCore::ARMEmitter::Reg::r11,
|
||||
FEXCore::ARMEmitter::Reg::r12,
|
||||
FEXCore::ARMEmitter::Reg::r13,
|
||||
FEXCore::ARMEmitter::Reg::r14,
|
||||
FEXCore::ARMEmitter::Reg::r15,
|
||||
FEXCore::ARMEmitter::Reg::r16,
|
||||
FEXCore::ARMEmitter::Reg::r17,
|
||||
FEXCore::ARMEmitter::Reg::r19,
|
||||
FEXCore::ARMEmitter::Reg::r29,
|
||||
constexpr std::array<ARMEmitter::Register, 18> SRA = {
|
||||
ARMEmitter::Reg::r4,
|
||||
ARMEmitter::Reg::r5,
|
||||
ARMEmitter::Reg::r6,
|
||||
ARMEmitter::Reg::r7,
|
||||
ARMEmitter::Reg::r8,
|
||||
ARMEmitter::Reg::r9,
|
||||
ARMEmitter::Reg::r10,
|
||||
ARMEmitter::Reg::r11,
|
||||
ARMEmitter::Reg::r12,
|
||||
ARMEmitter::Reg::r13,
|
||||
ARMEmitter::Reg::r14,
|
||||
ARMEmitter::Reg::r15,
|
||||
ARMEmitter::Reg::r16,
|
||||
ARMEmitter::Reg::r17,
|
||||
ARMEmitter::Reg::r19,
|
||||
ARMEmitter::Reg::r29,
|
||||
// PF/AF must be last.
|
||||
REG_PF,
|
||||
REG_AF,
|
||||
};
|
||||
|
||||
constexpr std::array<FEXCore::ARMEmitter::Register, 7> RA = {
|
||||
constexpr std::array<ARMEmitter::Register, 8> RA = {
|
||||
// All these callee saved
|
||||
FEXCore::ARMEmitter::Reg::r20, FEXCore::ARMEmitter::Reg::r21, FEXCore::ARMEmitter::Reg::r22, FEXCore::ARMEmitter::Reg::r23,
|
||||
FEXCore::ARMEmitter::Reg::r24, FEXCore::ARMEmitter::Reg::r25, FEXCore::ARMEmitter::Reg::r30,
|
||||
ARMEmitter::Reg::r20, ARMEmitter::Reg::r21, ARMEmitter::Reg::r22, ARMEmitter::Reg::r23,
|
||||
ARMEmitter::Reg::r24, ARMEmitter::Reg::r25, ARMEmitter::Reg::r30, ARMEmitter::Reg::r18,
|
||||
};
|
||||
|
||||
constexpr std::array<std::pair<FEXCore::ARMEmitter::Register, FEXCore::ARMEmitter::Register>, 3> RAPair = {{
|
||||
{FEXCore::ARMEmitter::Reg::r20, FEXCore::ARMEmitter::Reg::r21},
|
||||
{FEXCore::ARMEmitter::Reg::r22, FEXCore::ARMEmitter::Reg::r23},
|
||||
{FEXCore::ARMEmitter::Reg::r24, FEXCore::ARMEmitter::Reg::r25},
|
||||
}};
|
||||
constexpr unsigned RAPairs = 6;
|
||||
|
||||
// All are caller saved
|
||||
constexpr std::array<FEXCore::ARMEmitter::VRegister, 16> SRAFPR = {
|
||||
FEXCore::ARMEmitter::VReg::v16, FEXCore::ARMEmitter::VReg::v17, FEXCore::ARMEmitter::VReg::v18, FEXCore::ARMEmitter::VReg::v19,
|
||||
FEXCore::ARMEmitter::VReg::v20, FEXCore::ARMEmitter::VReg::v21, FEXCore::ARMEmitter::VReg::v22, FEXCore::ARMEmitter::VReg::v23,
|
||||
FEXCore::ARMEmitter::VReg::v24, FEXCore::ARMEmitter::VReg::v25, FEXCore::ARMEmitter::VReg::v26, FEXCore::ARMEmitter::VReg::v27,
|
||||
FEXCore::ARMEmitter::VReg::v28, FEXCore::ARMEmitter::VReg::v29, FEXCore::ARMEmitter::VReg::v30, FEXCore::ARMEmitter::VReg::v31};
|
||||
constexpr std::array<ARMEmitter::VRegister, 16> SRAFPR = {
|
||||
ARMEmitter::VReg::v16, ARMEmitter::VReg::v17, ARMEmitter::VReg::v18, ARMEmitter::VReg::v19,
|
||||
ARMEmitter::VReg::v20, ARMEmitter::VReg::v21, ARMEmitter::VReg::v22, ARMEmitter::VReg::v23,
|
||||
ARMEmitter::VReg::v24, ARMEmitter::VReg::v25, ARMEmitter::VReg::v26, ARMEmitter::VReg::v27,
|
||||
ARMEmitter::VReg::v28, ARMEmitter::VReg::v29, ARMEmitter::VReg::v30, ARMEmitter::VReg::v31};
|
||||
|
||||
// v8..v15 = (lower 64bits) Callee saved
|
||||
constexpr std::array<FEXCore::ARMEmitter::VRegister, 14> RAFPR = {
|
||||
constexpr std::array<ARMEmitter::VRegister, 14> RAFPR = {
|
||||
// v0 ~ v1 are used as temps.
|
||||
// FEXCore::ARMEmitter::VReg::v0, FEXCore::ARMEmitter::VReg::v1,
|
||||
// ARMEmitter::VReg::v0, ARMEmitter::VReg::v1,
|
||||
|
||||
FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3, FEXCore::ARMEmitter::VReg::v4, FEXCore::ARMEmitter::VReg::v5,
|
||||
FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7, FEXCore::ARMEmitter::VReg::v8, FEXCore::ARMEmitter::VReg::v9,
|
||||
FEXCore::ARMEmitter::VReg::v10, FEXCore::ARMEmitter::VReg::v11, FEXCore::ARMEmitter::VReg::v12, FEXCore::ARMEmitter::VReg::v13,
|
||||
FEXCore::ARMEmitter::VReg::v14, FEXCore::ARMEmitter::VReg::v15,
|
||||
ARMEmitter::VReg::v2, ARMEmitter::VReg::v3, ARMEmitter::VReg::v4, ARMEmitter::VReg::v5, ARMEmitter::VReg::v6,
|
||||
ARMEmitter::VReg::v7, ARMEmitter::VReg::v8, ARMEmitter::VReg::v9, ARMEmitter::VReg::v10, ARMEmitter::VReg::v11,
|
||||
ARMEmitter::VReg::v12, ARMEmitter::VReg::v13, ARMEmitter::VReg::v14, ARMEmitter::VReg::v15,
|
||||
};
|
||||
#else
|
||||
constexpr std::array<FEXCore::ARMEmitter::Register, 18> SRA = {
|
||||
FEXCore::ARMEmitter::Reg::r8,
|
||||
FEXCore::ARMEmitter::Reg::r0,
|
||||
FEXCore::ARMEmitter::Reg::r1,
|
||||
FEXCore::ARMEmitter::Reg::r27,
|
||||
constexpr std::array<ARMEmitter::Register, 18> SRA = {
|
||||
ARMEmitter::Reg::r8,
|
||||
ARMEmitter::Reg::r0,
|
||||
ARMEmitter::Reg::r1,
|
||||
ARMEmitter::Reg::r27,
|
||||
// SP's register location isn't specified by the ARM64EC ABI, we choose to use r23
|
||||
FEXCore::ARMEmitter::Reg::r23,
|
||||
FEXCore::ARMEmitter::Reg::r29,
|
||||
FEXCore::ARMEmitter::Reg::r25,
|
||||
FEXCore::ARMEmitter::Reg::r26,
|
||||
FEXCore::ARMEmitter::Reg::r2,
|
||||
FEXCore::ARMEmitter::Reg::r3,
|
||||
FEXCore::ARMEmitter::Reg::r4,
|
||||
FEXCore::ARMEmitter::Reg::r5,
|
||||
FEXCore::ARMEmitter::Reg::r19,
|
||||
FEXCore::ARMEmitter::Reg::r20,
|
||||
FEXCore::ARMEmitter::Reg::r21,
|
||||
FEXCore::ARMEmitter::Reg::r22,
|
||||
ARMEmitter::Reg::r23,
|
||||
ARMEmitter::Reg::r29,
|
||||
ARMEmitter::Reg::r25,
|
||||
ARMEmitter::Reg::r26,
|
||||
ARMEmitter::Reg::r2,
|
||||
ARMEmitter::Reg::r3,
|
||||
ARMEmitter::Reg::r4,
|
||||
ARMEmitter::Reg::r5,
|
||||
ARMEmitter::Reg::r19,
|
||||
ARMEmitter::Reg::r20,
|
||||
ARMEmitter::Reg::r21,
|
||||
ARMEmitter::Reg::r22,
|
||||
REG_PF,
|
||||
REG_AF,
|
||||
};
|
||||
|
||||
constexpr std::array<FEXCore::ARMEmitter::Register, 7> RA = {
|
||||
FEXCore::ARMEmitter::Reg::r6, FEXCore::ARMEmitter::Reg::r7, FEXCore::ARMEmitter::Reg::r14, FEXCore::ARMEmitter::Reg::r15,
|
||||
FEXCore::ARMEmitter::Reg::r16, FEXCore::ARMEmitter::Reg::r17, FEXCore::ARMEmitter::Reg::r30,
|
||||
constexpr std::array<ARMEmitter::Register, 7> RA = {
|
||||
ARMEmitter::Reg::r6, ARMEmitter::Reg::r7, ARMEmitter::Reg::r14, ARMEmitter::Reg::r15,
|
||||
ARMEmitter::Reg::r16, ARMEmitter::Reg::r17, ARMEmitter::Reg::r30,
|
||||
};
|
||||
|
||||
constexpr std::array<std::pair<FEXCore::ARMEmitter::Register, FEXCore::ARMEmitter::Register>, 3> RAPair = {{
|
||||
{FEXCore::ARMEmitter::Reg::r6, FEXCore::ARMEmitter::Reg::r7},
|
||||
{FEXCore::ARMEmitter::Reg::r14, FEXCore::ARMEmitter::Reg::r15},
|
||||
{FEXCore::ARMEmitter::Reg::r16, FEXCore::ARMEmitter::Reg::r17},
|
||||
}};
|
||||
constexpr unsigned RAPairs = 6;
|
||||
|
||||
constexpr std::array<FEXCore::ARMEmitter::VRegister, 16> SRAFPR = {
|
||||
FEXCore::ARMEmitter::VReg::v0, FEXCore::ARMEmitter::VReg::v1, FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3,
|
||||
FEXCore::ARMEmitter::VReg::v4, FEXCore::ARMEmitter::VReg::v5, FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7,
|
||||
FEXCore::ARMEmitter::VReg::v8, FEXCore::ARMEmitter::VReg::v9, FEXCore::ARMEmitter::VReg::v10, FEXCore::ARMEmitter::VReg::v11,
|
||||
FEXCore::ARMEmitter::VReg::v12, FEXCore::ARMEmitter::VReg::v13, FEXCore::ARMEmitter::VReg::v14, FEXCore::ARMEmitter::VReg::v15,
|
||||
constexpr std::array<ARMEmitter::VRegister, 16> SRAFPR = {
|
||||
ARMEmitter::VReg::v0, ARMEmitter::VReg::v1, ARMEmitter::VReg::v2, ARMEmitter::VReg::v3,
|
||||
ARMEmitter::VReg::v4, ARMEmitter::VReg::v5, ARMEmitter::VReg::v6, ARMEmitter::VReg::v7,
|
||||
ARMEmitter::VReg::v8, ARMEmitter::VReg::v9, ARMEmitter::VReg::v10, ARMEmitter::VReg::v11,
|
||||
ARMEmitter::VReg::v12, ARMEmitter::VReg::v13, ARMEmitter::VReg::v14, ARMEmitter::VReg::v15,
|
||||
};
|
||||
|
||||
constexpr std::array<FEXCore::ARMEmitter::VRegister, 14> RAFPR = {
|
||||
FEXCore::ARMEmitter::VReg::v18, FEXCore::ARMEmitter::VReg::v19, FEXCore::ARMEmitter::VReg::v20, FEXCore::ARMEmitter::VReg::v21,
|
||||
FEXCore::ARMEmitter::VReg::v22, FEXCore::ARMEmitter::VReg::v23, FEXCore::ARMEmitter::VReg::v24, FEXCore::ARMEmitter::VReg::v25,
|
||||
FEXCore::ARMEmitter::VReg::v26, FEXCore::ARMEmitter::VReg::v27, FEXCore::ARMEmitter::VReg::v28, FEXCore::ARMEmitter::VReg::v29,
|
||||
FEXCore::ARMEmitter::VReg::v30, FEXCore::ARMEmitter::VReg::v31};
|
||||
constexpr std::array<ARMEmitter::VRegister, 14> RAFPR = {
|
||||
ARMEmitter::VReg::v18, ARMEmitter::VReg::v19, ARMEmitter::VReg::v20, ARMEmitter::VReg::v21, ARMEmitter::VReg::v22,
|
||||
ARMEmitter::VReg::v23, ARMEmitter::VReg::v24, ARMEmitter::VReg::v25, ARMEmitter::VReg::v26, ARMEmitter::VReg::v27,
|
||||
ARMEmitter::VReg::v28, ARMEmitter::VReg::v29, ARMEmitter::VReg::v30, ARMEmitter::VReg::v31};
|
||||
#endif
|
||||
|
||||
// I wish this could get constexpr generated from SRA's definition but impossible until libstdc++12, libc++15.
|
||||
// SRA GPRs that need to be spilled when calling a function with `preserve_all` ABI.
|
||||
constexpr std::array<FEXCore::ARMEmitter::Register, 7> PreserveAll_SRA = {
|
||||
FEXCore::ARMEmitter::Reg::r4, FEXCore::ARMEmitter::Reg::r5, FEXCore::ARMEmitter::Reg::r6, FEXCore::ARMEmitter::Reg::r7,
|
||||
FEXCore::ARMEmitter::Reg::r8, FEXCore::ARMEmitter::Reg::r16, FEXCore::ARMEmitter::Reg::r17,
|
||||
constexpr std::array<ARMEmitter::Register, 7> PreserveAll_SRA = {
|
||||
ARMEmitter::Reg::r4, ARMEmitter::Reg::r5, ARMEmitter::Reg::r6, ARMEmitter::Reg::r7,
|
||||
ARMEmitter::Reg::r8, ARMEmitter::Reg::r16, ARMEmitter::Reg::r17,
|
||||
};
|
||||
|
||||
constexpr uint32_t PreserveAll_SRAMask = {[]() -> uint32_t {
|
||||
@@ -160,12 +146,12 @@ namespace x64 {
|
||||
}()};
|
||||
|
||||
// Dynamic GPRs
|
||||
constexpr std::array<FEXCore::ARMEmitter::Register, 1> PreserveAll_Dynamic = {
|
||||
constexpr std::array<ARMEmitter::Register, 1> PreserveAll_Dynamic = {
|
||||
// Only LR needs to get saved.
|
||||
FEXCore::ARMEmitter::Reg::r30};
|
||||
ARMEmitter::Reg::r30};
|
||||
|
||||
// SRA FPRs that need to be spilled when calling a function with `preserve_all` ABI.
|
||||
constexpr std::array<FEXCore::ARMEmitter::Register, 0> PreserveAll_SRAFPR = {
|
||||
constexpr std::array<ARMEmitter::Register, 0> PreserveAll_SRAFPR = {
|
||||
// None.
|
||||
};
|
||||
|
||||
@@ -179,15 +165,14 @@ namespace x64 {
|
||||
|
||||
// Dynamic FPRs
|
||||
// - v0-v7
|
||||
constexpr std::array<FEXCore::ARMEmitter::VRegister, 6> PreserveAll_DynamicFPR = {
|
||||
constexpr std::array<ARMEmitter::VRegister, 6> PreserveAll_DynamicFPR = {
|
||||
// v0 ~ v1 are temps
|
||||
FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3, FEXCore::ARMEmitter::VReg::v4,
|
||||
FEXCore::ARMEmitter::VReg::v5, FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7,
|
||||
ARMEmitter::VReg::v2, ARMEmitter::VReg::v3, ARMEmitter::VReg::v4, ARMEmitter::VReg::v5, ARMEmitter::VReg::v6, ARMEmitter::VReg::v7,
|
||||
};
|
||||
|
||||
// SRA FPRs that need to be spilled when the host supports SVE-256bit with `preserve_all` ABI.
|
||||
// This is /all/ of the SRA registers
|
||||
constexpr std::array<FEXCore::ARMEmitter::VRegister, 16> PreserveAll_SRAFPRSVE = SRAFPR;
|
||||
constexpr std::array<ARMEmitter::VRegister, 16> PreserveAll_SRAFPRSVE = SRAFPR;
|
||||
|
||||
constexpr uint32_t PreserveAll_SRAFPRSVEMask = {[]() -> uint32_t {
|
||||
uint32_t Mask {};
|
||||
@@ -198,89 +183,77 @@ namespace x64 {
|
||||
}()};
|
||||
|
||||
// Dynamic FPRs when the host supports SVE-256bit.
|
||||
constexpr std::array<FEXCore::ARMEmitter::VRegister, 14> PreserveAll_DynamicFPRSVE = {
|
||||
constexpr std::array<ARMEmitter::VRegister, 14> PreserveAll_DynamicFPRSVE = {
|
||||
// v0 ~ v1 are used as temps.
|
||||
FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3, FEXCore::ARMEmitter::VReg::v4, FEXCore::ARMEmitter::VReg::v5,
|
||||
FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7, FEXCore::ARMEmitter::VReg::v8, FEXCore::ARMEmitter::VReg::v9,
|
||||
FEXCore::ARMEmitter::VReg::v10, FEXCore::ARMEmitter::VReg::v11, FEXCore::ARMEmitter::VReg::v12, FEXCore::ARMEmitter::VReg::v13,
|
||||
FEXCore::ARMEmitter::VReg::v14, FEXCore::ARMEmitter::VReg::v15,
|
||||
ARMEmitter::VReg::v2, ARMEmitter::VReg::v3, ARMEmitter::VReg::v4, ARMEmitter::VReg::v5, ARMEmitter::VReg::v6,
|
||||
ARMEmitter::VReg::v7, ARMEmitter::VReg::v8, ARMEmitter::VReg::v9, ARMEmitter::VReg::v10, ARMEmitter::VReg::v11,
|
||||
ARMEmitter::VReg::v12, ARMEmitter::VReg::v13, ARMEmitter::VReg::v14, ARMEmitter::VReg::v15,
|
||||
};
|
||||
} // namespace x64
|
||||
|
||||
namespace x32 {
|
||||
// All but x19 and x29 are caller saved
|
||||
constexpr std::array<FEXCore::ARMEmitter::Register, 10> SRA = {
|
||||
FEXCore::ARMEmitter::Reg::r4,
|
||||
FEXCore::ARMEmitter::Reg::r5,
|
||||
FEXCore::ARMEmitter::Reg::r6,
|
||||
FEXCore::ARMEmitter::Reg::r7,
|
||||
FEXCore::ARMEmitter::Reg::r8,
|
||||
FEXCore::ARMEmitter::Reg::r9,
|
||||
FEXCore::ARMEmitter::Reg::r10,
|
||||
FEXCore::ARMEmitter::Reg::r11,
|
||||
constexpr std::array<ARMEmitter::Register, 10> SRA = {
|
||||
ARMEmitter::Reg::r4,
|
||||
ARMEmitter::Reg::r5,
|
||||
ARMEmitter::Reg::r6,
|
||||
ARMEmitter::Reg::r7,
|
||||
ARMEmitter::Reg::r8,
|
||||
ARMEmitter::Reg::r9,
|
||||
ARMEmitter::Reg::r10,
|
||||
ARMEmitter::Reg::r11,
|
||||
// PF/AF must be last.
|
||||
REG_PF,
|
||||
REG_AF,
|
||||
};
|
||||
|
||||
constexpr std::array<FEXCore::ARMEmitter::Register, 15> RA = {
|
||||
constexpr std::array<ARMEmitter::Register, 15> RA = {
|
||||
// All these callee saved
|
||||
FEXCore::ARMEmitter::Reg::r20,
|
||||
FEXCore::ARMEmitter::Reg::r21,
|
||||
FEXCore::ARMEmitter::Reg::r22,
|
||||
FEXCore::ARMEmitter::Reg::r23,
|
||||
FEXCore::ARMEmitter::Reg::r24,
|
||||
FEXCore::ARMEmitter::Reg::r25,
|
||||
ARMEmitter::Reg::r20,
|
||||
ARMEmitter::Reg::r21,
|
||||
ARMEmitter::Reg::r22,
|
||||
ARMEmitter::Reg::r23,
|
||||
ARMEmitter::Reg::r24,
|
||||
ARMEmitter::Reg::r25,
|
||||
|
||||
// Registers only available on 32-bit
|
||||
// All these are caller saved (except for r19).
|
||||
FEXCore::ARMEmitter::Reg::r12,
|
||||
FEXCore::ARMEmitter::Reg::r13,
|
||||
FEXCore::ARMEmitter::Reg::r14,
|
||||
FEXCore::ARMEmitter::Reg::r15,
|
||||
FEXCore::ARMEmitter::Reg::r16,
|
||||
FEXCore::ARMEmitter::Reg::r17,
|
||||
FEXCore::ARMEmitter::Reg::r29,
|
||||
FEXCore::ARMEmitter::Reg::r30,
|
||||
ARMEmitter::Reg::r12,
|
||||
ARMEmitter::Reg::r13,
|
||||
ARMEmitter::Reg::r14,
|
||||
ARMEmitter::Reg::r15,
|
||||
ARMEmitter::Reg::r16,
|
||||
ARMEmitter::Reg::r17,
|
||||
ARMEmitter::Reg::r29,
|
||||
ARMEmitter::Reg::r30,
|
||||
|
||||
FEXCore::ARMEmitter::Reg::r19,
|
||||
ARMEmitter::Reg::r19,
|
||||
};
|
||||
|
||||
constexpr std::array<std::pair<FEXCore::ARMEmitter::Register, FEXCore::ARMEmitter::Register>, 7> RAPair = {{
|
||||
{FEXCore::ARMEmitter::Reg::r20, FEXCore::ARMEmitter::Reg::r21},
|
||||
{FEXCore::ARMEmitter::Reg::r22, FEXCore::ARMEmitter::Reg::r23},
|
||||
{FEXCore::ARMEmitter::Reg::r24, FEXCore::ARMEmitter::Reg::r25},
|
||||
|
||||
{FEXCore::ARMEmitter::Reg::r12, FEXCore::ARMEmitter::Reg::r13},
|
||||
{FEXCore::ARMEmitter::Reg::r14, FEXCore::ARMEmitter::Reg::r15},
|
||||
{FEXCore::ARMEmitter::Reg::r16, FEXCore::ARMEmitter::Reg::r17},
|
||||
{FEXCore::ARMEmitter::Reg::r29, FEXCore::ARMEmitter::Reg::r30},
|
||||
}};
|
||||
constexpr unsigned RAPairs = 12;
|
||||
|
||||
// All are caller saved
|
||||
constexpr std::array<FEXCore::ARMEmitter::VRegister, 8> SRAFPR = {
|
||||
FEXCore::ARMEmitter::VReg::v16, FEXCore::ARMEmitter::VReg::v17, FEXCore::ARMEmitter::VReg::v18, FEXCore::ARMEmitter::VReg::v19,
|
||||
FEXCore::ARMEmitter::VReg::v20, FEXCore::ARMEmitter::VReg::v21, FEXCore::ARMEmitter::VReg::v22, FEXCore::ARMEmitter::VReg::v23,
|
||||
constexpr std::array<ARMEmitter::VRegister, 8> SRAFPR = {
|
||||
ARMEmitter::VReg::v16, ARMEmitter::VReg::v17, ARMEmitter::VReg::v18, ARMEmitter::VReg::v19,
|
||||
ARMEmitter::VReg::v20, ARMEmitter::VReg::v21, ARMEmitter::VReg::v22, ARMEmitter::VReg::v23,
|
||||
};
|
||||
|
||||
// v8..v15 = (lower 64bits) Callee saved
|
||||
constexpr std::array<FEXCore::ARMEmitter::VRegister, 22> RAFPR = {
|
||||
constexpr std::array<ARMEmitter::VRegister, 22> RAFPR = {
|
||||
// v0 ~ v1 are used as temps.
|
||||
// FEXCore::ARMEmitter::VReg::v0, FEXCore::ARMEmitter::VReg::v1,
|
||||
// ARMEmitter::VReg::v0, ARMEmitter::VReg::v1,
|
||||
|
||||
FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3, FEXCore::ARMEmitter::VReg::v4, FEXCore::ARMEmitter::VReg::v5,
|
||||
FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7, FEXCore::ARMEmitter::VReg::v8, FEXCore::ARMEmitter::VReg::v9,
|
||||
FEXCore::ARMEmitter::VReg::v10, FEXCore::ARMEmitter::VReg::v11, FEXCore::ARMEmitter::VReg::v12, FEXCore::ARMEmitter::VReg::v13,
|
||||
FEXCore::ARMEmitter::VReg::v14, FEXCore::ARMEmitter::VReg::v15,
|
||||
ARMEmitter::VReg::v2, ARMEmitter::VReg::v3, ARMEmitter::VReg::v4, ARMEmitter::VReg::v5, ARMEmitter::VReg::v6,
|
||||
ARMEmitter::VReg::v7, ARMEmitter::VReg::v8, ARMEmitter::VReg::v9, ARMEmitter::VReg::v10, ARMEmitter::VReg::v11,
|
||||
ARMEmitter::VReg::v12, ARMEmitter::VReg::v13, ARMEmitter::VReg::v14, ARMEmitter::VReg::v15,
|
||||
|
||||
FEXCore::ARMEmitter::VReg::v24, FEXCore::ARMEmitter::VReg::v25, FEXCore::ARMEmitter::VReg::v26, FEXCore::ARMEmitter::VReg::v27,
|
||||
FEXCore::ARMEmitter::VReg::v28, FEXCore::ARMEmitter::VReg::v29, FEXCore::ARMEmitter::VReg::v30, FEXCore::ARMEmitter::VReg::v31};
|
||||
ARMEmitter::VReg::v24, ARMEmitter::VReg::v25, ARMEmitter::VReg::v26, ARMEmitter::VReg::v27, ARMEmitter::VReg::v28,
|
||||
ARMEmitter::VReg::v29, ARMEmitter::VReg::v30, ARMEmitter::VReg::v31};
|
||||
|
||||
// I wish this could get constexpr generated from SRA's definition but impossible until libstdc++12, libc++15.
|
||||
// SRA GPRs that need to be spilled when calling a function with `preserve_all` ABI.
|
||||
constexpr std::array<FEXCore::ARMEmitter::Register, 5> PreserveAll_SRA = {
|
||||
FEXCore::ARMEmitter::Reg::r4, FEXCore::ARMEmitter::Reg::r5, FEXCore::ARMEmitter::Reg::r6,
|
||||
FEXCore::ARMEmitter::Reg::r7, FEXCore::ARMEmitter::Reg::r8,
|
||||
constexpr std::array<ARMEmitter::Register, 5> PreserveAll_SRA = {
|
||||
ARMEmitter::Reg::r4, ARMEmitter::Reg::r5, ARMEmitter::Reg::r6, ARMEmitter::Reg::r7, ARMEmitter::Reg::r8,
|
||||
};
|
||||
|
||||
constexpr uint32_t PreserveAll_SRAMask = {[]() -> uint32_t {
|
||||
@@ -306,11 +279,10 @@ namespace x32 {
|
||||
}()};
|
||||
|
||||
// Dynamic GPRs
|
||||
constexpr std::array<FEXCore::ARMEmitter::Register, 3> PreserveAll_Dynamic = {
|
||||
FEXCore::ARMEmitter::Reg::r16, FEXCore::ARMEmitter::Reg::r17, FEXCore::ARMEmitter::Reg::r30};
|
||||
constexpr std::array<ARMEmitter::Register, 3> PreserveAll_Dynamic = {ARMEmitter::Reg::r16, ARMEmitter::Reg::r17, ARMEmitter::Reg::r30};
|
||||
|
||||
// SRA FPRs that need to be spilled when calling a function with `preserve_all` ABI.
|
||||
constexpr std::array<FEXCore::ARMEmitter::Register, 0> PreserveAll_SRAFPR = {
|
||||
constexpr std::array<ARMEmitter::Register, 0> PreserveAll_SRAFPR = {
|
||||
// None.
|
||||
};
|
||||
|
||||
@@ -324,15 +296,14 @@ namespace x32 {
|
||||
|
||||
// Dynamic FPRs
|
||||
// - v0-v7
|
||||
constexpr std::array<FEXCore::ARMEmitter::VRegister, 6> PreserveAll_DynamicFPR = {
|
||||
constexpr std::array<ARMEmitter::VRegister, 6> PreserveAll_DynamicFPR = {
|
||||
// v0 ~ v1 are temps
|
||||
FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3, FEXCore::ARMEmitter::VReg::v4,
|
||||
FEXCore::ARMEmitter::VReg::v5, FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7,
|
||||
ARMEmitter::VReg::v2, ARMEmitter::VReg::v3, ARMEmitter::VReg::v4, ARMEmitter::VReg::v5, ARMEmitter::VReg::v6, ARMEmitter::VReg::v7,
|
||||
};
|
||||
|
||||
// SRA FPRs that need to be spilled when the host supports SVE-256bit with `preserve_all` ABI.
|
||||
// This is /all/ of the SRA registers
|
||||
constexpr std::array<FEXCore::ARMEmitter::VRegister, 8> PreserveAll_SRAFPRSVE = SRAFPR;
|
||||
constexpr std::array<ARMEmitter::VRegister, 8> PreserveAll_SRAFPRSVE = SRAFPR;
|
||||
|
||||
constexpr uint32_t PreserveAll_SRAFPRSVEMask = {[]() -> uint32_t {
|
||||
uint32_t Mask {};
|
||||
@@ -343,15 +314,14 @@ namespace x32 {
|
||||
}()};
|
||||
|
||||
// Dynamic FPRs when the host supports SVE-256bit.
|
||||
constexpr std::array<FEXCore::ARMEmitter::VRegister, 22> PreserveAll_DynamicFPRSVE = {
|
||||
constexpr std::array<ARMEmitter::VRegister, 22> PreserveAll_DynamicFPRSVE = {
|
||||
// v0 ~ v1 are used as temps.
|
||||
FEXCore::ARMEmitter::VReg::v2, FEXCore::ARMEmitter::VReg::v3, FEXCore::ARMEmitter::VReg::v4, FEXCore::ARMEmitter::VReg::v5,
|
||||
FEXCore::ARMEmitter::VReg::v6, FEXCore::ARMEmitter::VReg::v7, FEXCore::ARMEmitter::VReg::v8, FEXCore::ARMEmitter::VReg::v9,
|
||||
FEXCore::ARMEmitter::VReg::v10, FEXCore::ARMEmitter::VReg::v11, FEXCore::ARMEmitter::VReg::v12, FEXCore::ARMEmitter::VReg::v13,
|
||||
FEXCore::ARMEmitter::VReg::v14, FEXCore::ARMEmitter::VReg::v15,
|
||||
ARMEmitter::VReg::v2, ARMEmitter::VReg::v3, ARMEmitter::VReg::v4, ARMEmitter::VReg::v5, ARMEmitter::VReg::v6,
|
||||
ARMEmitter::VReg::v7, ARMEmitter::VReg::v8, ARMEmitter::VReg::v9, ARMEmitter::VReg::v10, ARMEmitter::VReg::v11,
|
||||
ARMEmitter::VReg::v12, ARMEmitter::VReg::v13, ARMEmitter::VReg::v14, ARMEmitter::VReg::v15,
|
||||
|
||||
FEXCore::ARMEmitter::VReg::v24, FEXCore::ARMEmitter::VReg::v25, FEXCore::ARMEmitter::VReg::v26, FEXCore::ARMEmitter::VReg::v27,
|
||||
FEXCore::ARMEmitter::VReg::v28, FEXCore::ARMEmitter::VReg::v29, FEXCore::ARMEmitter::VReg::v30, FEXCore::ARMEmitter::VReg::v31};
|
||||
ARMEmitter::VReg::v24, ARMEmitter::VReg::v25, ARMEmitter::VReg::v26, ARMEmitter::VReg::v27, ARMEmitter::VReg::v28,
|
||||
ARMEmitter::VReg::v29, ARMEmitter::VReg::v30, ARMEmitter::VReg::v31};
|
||||
} // namespace x32
|
||||
|
||||
// We want vixl to not allocate a default buffer. Jit and dispatcher will manually create one.
|
||||
@@ -385,18 +355,18 @@ Arm64Emitter::Arm64Emitter(FEXCore::Context::ContextImpl* ctx, void* EmissionPtr
|
||||
if (EmitterCTX->Config.Is64BitMode()) {
|
||||
StaticRegisters = x64::SRA;
|
||||
GeneralRegisters = x64::RA;
|
||||
GeneralPairRegisters = x64::RAPair;
|
||||
StaticFPRegisters = x64::SRAFPR;
|
||||
GeneralFPRegisters = x64::RAFPR;
|
||||
PairRegisters = x64::RAPairs;
|
||||
#ifdef _M_ARM_64EC
|
||||
ConfiguredDynamicRegisterBase = std::span(x64::RA.begin(), 7);
|
||||
#endif
|
||||
} else {
|
||||
ConfiguredDynamicRegisterBase = std::span(x32::RA.begin() + 6, 8);
|
||||
PairRegisters = x32::RAPairs;
|
||||
|
||||
StaticRegisters = x32::SRA;
|
||||
GeneralRegisters = x32::RA;
|
||||
GeneralPairRegisters = x32::RAPair;
|
||||
|
||||
StaticFPRegisters = x32::SRAFPR;
|
||||
GeneralFPRegisters = x32::RAFPR;
|
||||
@@ -463,6 +433,17 @@ void Arm64Emitter::LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, ui
|
||||
}
|
||||
}
|
||||
|
||||
// If we can't handle negatives with the orr, try with movn+movk
|
||||
if (Is64Bit && ((~Constant) >> 32) == 0) {
|
||||
movn(s, Reg, (~Constant) & 0xFFFF);
|
||||
movk(s, Reg, (Constant >> 16) & 0xFFFF, 16);
|
||||
if (NOPPad) {
|
||||
nop();
|
||||
nop();
|
||||
}
|
||||
return;
|
||||
}
|
||||
|
||||
// ADRP+ADD is specifically optimized in hardware
|
||||
// Check if we can use this
|
||||
auto PC = GetCursorAddress<uint64_t>();
|
||||
@@ -589,7 +570,7 @@ void Arm64Emitter::PopCalleeSavedRegisters() {
|
||||
}
|
||||
}
|
||||
|
||||
void Arm64Emitter::SpillStaticRegs(FEXCore::ARMEmitter::Register TmpReg, bool FPRs, uint32_t GPRSpillMask, uint32_t FPRSpillMask) {
|
||||
void Arm64Emitter::SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs, uint32_t GPRSpillMask, uint32_t FPRSpillMask) {
|
||||
#ifndef VIXL_SIMULATOR
|
||||
if (EmitterCTX->HostFeatures.SupportsAFP) {
|
||||
// Disable AFP features when spilling registers.
|
||||
@@ -643,7 +624,7 @@ void Arm64Emitter::SpillStaticRegs(FEXCore::ARMEmitter::Register TmpReg, bool FP
|
||||
}
|
||||
|
||||
if (FPRs) {
|
||||
if (EmitterCTX->HostFeatures.SupportsAVX) {
|
||||
if (EmitterCTX->HostFeatures.SupportsAVX && EmitterCTX->HostFeatures.SupportsSVE256) {
|
||||
for (size_t i = 0; i < StaticFPRegisters.size(); i++) {
|
||||
const auto Reg = StaticFPRegisters[i];
|
||||
|
||||
@@ -683,7 +664,7 @@ void Arm64Emitter::SpillStaticRegs(FEXCore::ARMEmitter::Register TmpReg, bool FP
|
||||
}
|
||||
|
||||
void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRFillMask) {
|
||||
FEXCore::ARMEmitter::Register TmpReg = FEXCore::ARMEmitter::Reg::r0;
|
||||
ARMEmitter::Register TmpReg = ARMEmitter::Reg::r0;
|
||||
LOGMAN_THROW_A_FMT(GPRFillMask != 0, "Must fill at least 1 GPR for a temp");
|
||||
[[maybe_unused]] bool FoundRegister {};
|
||||
for (auto Reg : StaticRegisters) {
|
||||
@@ -728,13 +709,15 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
|
||||
// We don't bother spilling these in SpillStaticRegs,
|
||||
// since all that matters is we restore them on a fill.
|
||||
// It's not a concern if they get trounced by something else.
|
||||
if (EmitterCTX->HostFeatures.SupportsSVE) {
|
||||
if (EmitterCTX->HostFeatures.SupportsSVE128) {
|
||||
ptrue(ARMEmitter::SubRegSize::i8Bit, PRED_TMP_16B, ARMEmitter::PredicatePattern::SVE_VL16);
|
||||
}
|
||||
|
||||
if (EmitterCTX->HostFeatures.SupportsAVX) {
|
||||
if (EmitterCTX->HostFeatures.SupportsSVE256) {
|
||||
ptrue(ARMEmitter::SubRegSize::i8Bit, PRED_TMP_32B, ARMEmitter::PredicatePattern::SVE_VL32);
|
||||
}
|
||||
|
||||
if (EmitterCTX->HostFeatures.SupportsAVX && EmitterCTX->HostFeatures.SupportsSVE256) {
|
||||
for (size_t i = 0; i < StaticFPRegisters.size(); i++) {
|
||||
const auto Reg = StaticFPRegisters[i];
|
||||
if (((1U << Reg.Idx()) & FPRFillMask) != 0) {
|
||||
@@ -798,8 +781,8 @@ void Arm64Emitter::FillStaticRegs(bool FPRs, uint32_t GPRFillMask, uint32_t FPRF
|
||||
}
|
||||
}
|
||||
|
||||
void Arm64Emitter::PushVectorRegisters(FEXCore::ARMEmitter::Register TmpReg, bool SVERegs, std::span<const FEXCore::ARMEmitter::VRegister> VRegs) {
|
||||
if (SVERegs) {
|
||||
void Arm64Emitter::PushVectorRegisters(ARMEmitter::Register TmpReg, bool SVE256Regs, std::span<const ARMEmitter::VRegister> VRegs) {
|
||||
if (SVE256Regs) {
|
||||
size_t i = 0;
|
||||
|
||||
for (; i < (VRegs.size() % 4); i += 2) {
|
||||
@@ -835,7 +818,7 @@ void Arm64Emitter::PushVectorRegisters(FEXCore::ARMEmitter::Register TmpReg, boo
|
||||
}
|
||||
}
|
||||
|
||||
void Arm64Emitter::PushGeneralRegisters(FEXCore::ARMEmitter::Register TmpReg, std::span<const FEXCore::ARMEmitter::Register> Regs) {
|
||||
void Arm64Emitter::PushGeneralRegisters(ARMEmitter::Register TmpReg, std::span<const ARMEmitter::Register> Regs) {
|
||||
size_t i = 0;
|
||||
for (; i < (Regs.size() % 2); ++i) {
|
||||
const auto Reg1 = Regs[i];
|
||||
@@ -849,8 +832,8 @@ void Arm64Emitter::PushGeneralRegisters(FEXCore::ARMEmitter::Register TmpReg, st
|
||||
}
|
||||
}
|
||||
|
||||
void Arm64Emitter::PopVectorRegisters(bool SVERegs, std::span<const FEXCore::ARMEmitter::VRegister> VRegs) {
|
||||
if (SVERegs) {
|
||||
void Arm64Emitter::PopVectorRegisters(bool SVE256Regs, std::span<const ARMEmitter::VRegister> VRegs) {
|
||||
if (SVE256Regs) {
|
||||
size_t i = 0;
|
||||
for (; i < (VRegs.size() % 4); i += 2) {
|
||||
const auto Reg1 = VRegs[i];
|
||||
@@ -885,7 +868,7 @@ void Arm64Emitter::PopVectorRegisters(bool SVERegs, std::span<const FEXCore::ARM
|
||||
}
|
||||
}
|
||||
|
||||
void Arm64Emitter::PopGeneralRegisters(std::span<const FEXCore::ARMEmitter::Register> Regs) {
|
||||
void Arm64Emitter::PopGeneralRegisters(std::span<const ARMEmitter::Register> Regs) {
|
||||
size_t i = 0;
|
||||
for (; i < (Regs.size() % 2); ++i) {
|
||||
const auto Reg1 = Regs[i];
|
||||
@@ -898,10 +881,10 @@ void Arm64Emitter::PopGeneralRegisters(std::span<const FEXCore::ARMEmitter::Regi
|
||||
}
|
||||
}
|
||||
|
||||
void Arm64Emitter::PushDynamicRegsAndLR(FEXCore::ARMEmitter::Register TmpReg) {
|
||||
const auto CanUseSVE = EmitterCTX->HostFeatures.SupportsAVX;
|
||||
void Arm64Emitter::PushDynamicRegsAndLR(ARMEmitter::Register TmpReg) {
|
||||
const auto CanUseSVE256 = EmitterCTX->HostFeatures.SupportsSVE256;
|
||||
const auto GPRSize = (ConfiguredDynamicRegisterBase.size() + 1) * Core::CPUState::GPR_REG_SIZE;
|
||||
const auto FPRRegSize = CanUseSVE ? Core::CPUState::XMM_AVX_REG_SIZE : Core::CPUState::XMM_SSE_REG_SIZE;
|
||||
const auto FPRRegSize = CanUseSVE256 ? 32 : 16;
|
||||
const auto FPRSize = GeneralFPRegisters.size() * FPRRegSize;
|
||||
const uint64_t SPOffset = AlignUp(GPRSize + FPRSize, 16);
|
||||
|
||||
@@ -913,7 +896,7 @@ void Arm64Emitter::PushDynamicRegsAndLR(FEXCore::ARMEmitter::Register TmpReg) {
|
||||
LOGMAN_THROW_A_FMT(GeneralFPRegisters.size() % 2 == 0, "Needs to have multiple of 2 FPRs for RA");
|
||||
|
||||
// Push the vector registers
|
||||
PushVectorRegisters(TmpReg, CanUseSVE, GeneralFPRegisters);
|
||||
PushVectorRegisters(TmpReg, CanUseSVE256, GeneralFPRegisters);
|
||||
|
||||
// Push the general registers.
|
||||
PushGeneralRegisters(TmpReg, ConfiguredDynamicRegisterBase);
|
||||
@@ -924,10 +907,10 @@ void Arm64Emitter::PushDynamicRegsAndLR(FEXCore::ARMEmitter::Register TmpReg) {
|
||||
}
|
||||
|
||||
void Arm64Emitter::PopDynamicRegsAndLR() {
|
||||
const auto CanUseSVE = EmitterCTX->HostFeatures.SupportsAVX;
|
||||
const auto CanUseSVE256 = EmitterCTX->HostFeatures.SupportsSVE256;
|
||||
|
||||
// Pop vectors first
|
||||
PopVectorRegisters(CanUseSVE, GeneralFPRegisters);
|
||||
PopVectorRegisters(CanUseSVE256, GeneralFPRegisters);
|
||||
|
||||
// Pop GPRs second
|
||||
PopGeneralRegisters(ConfiguredDynamicRegisterBase);
|
||||
@@ -937,12 +920,12 @@ void Arm64Emitter::PopDynamicRegsAndLR() {
|
||||
#endif
|
||||
}
|
||||
|
||||
void Arm64Emitter::SpillForPreserveAllABICall(FEXCore::ARMEmitter::Register TmpReg, bool FPRs) {
|
||||
const auto CanUseSVE = EmitterCTX->HostFeatures.SupportsAVX;
|
||||
const auto FPRRegSize = CanUseSVE ? Core::CPUState::XMM_AVX_REG_SIZE : Core::CPUState::XMM_SSE_REG_SIZE;
|
||||
void Arm64Emitter::SpillForPreserveAllABICall(ARMEmitter::Register TmpReg, bool FPRs) {
|
||||
const auto CanUseSVE256 = EmitterCTX->HostFeatures.SupportsSVE256;
|
||||
const auto FPRRegSize = CanUseSVE256 ? 32 : 16;
|
||||
|
||||
std::span<const FEXCore::ARMEmitter::Register> DynamicGPRs {};
|
||||
std::span<const FEXCore::ARMEmitter::VRegister> DynamicFPRs {};
|
||||
std::span<const ARMEmitter::Register> DynamicGPRs {};
|
||||
std::span<const ARMEmitter::VRegister> DynamicFPRs {};
|
||||
uint32_t PreserveSRAMask {};
|
||||
uint32_t PreserveSRAFPRMask {};
|
||||
if (EmitterCTX->Config.Is64BitMode()) {
|
||||
@@ -951,7 +934,7 @@ void Arm64Emitter::SpillForPreserveAllABICall(FEXCore::ARMEmitter::Register TmpR
|
||||
PreserveSRAMask = x64::PreserveAll_SRAMask;
|
||||
PreserveSRAFPRMask = x64::PreserveAll_SRAFPRMask;
|
||||
|
||||
if (CanUseSVE) {
|
||||
if (CanUseSVE256) {
|
||||
DynamicFPRs = x64::PreserveAll_DynamicFPRSVE;
|
||||
PreserveSRAFPRMask = x64::PreserveAll_SRAFPRSVEMask;
|
||||
}
|
||||
@@ -961,7 +944,7 @@ void Arm64Emitter::SpillForPreserveAllABICall(FEXCore::ARMEmitter::Register TmpR
|
||||
PreserveSRAMask = x32::PreserveAll_SRAMask;
|
||||
PreserveSRAFPRMask = x32::PreserveAll_SRAFPRMask;
|
||||
|
||||
if (CanUseSVE) {
|
||||
if (CanUseSVE256) {
|
||||
DynamicFPRs = x32::PreserveAll_DynamicFPRSVE;
|
||||
PreserveSRAFPRMask = x32::PreserveAll_SRAFPRSVEMask;
|
||||
}
|
||||
@@ -980,17 +963,17 @@ void Arm64Emitter::SpillForPreserveAllABICall(FEXCore::ARMEmitter::Register TmpR
|
||||
add(ARMEmitter::Size::i64Bit, TmpReg, ARMEmitter::Reg::rsp, 0);
|
||||
|
||||
// Push the vector registers.
|
||||
PushVectorRegisters(TmpReg, CanUseSVE, DynamicFPRs);
|
||||
PushVectorRegisters(TmpReg, CanUseSVE256, DynamicFPRs);
|
||||
|
||||
// Push the general registers.
|
||||
PushGeneralRegisters(TmpReg, DynamicGPRs);
|
||||
}
|
||||
|
||||
void Arm64Emitter::FillForPreserveAllABICall(bool FPRs) {
|
||||
const auto CanUseSVE = EmitterCTX->HostFeatures.SupportsAVX;
|
||||
const auto CanUseSVE256 = EmitterCTX->HostFeatures.SupportsSVE256;
|
||||
|
||||
std::span<const FEXCore::ARMEmitter::Register> DynamicGPRs {};
|
||||
std::span<const FEXCore::ARMEmitter::VRegister> DynamicFPRs {};
|
||||
std::span<const ARMEmitter::Register> DynamicGPRs {};
|
||||
std::span<const ARMEmitter::VRegister> DynamicFPRs {};
|
||||
uint32_t PreserveSRAMask {};
|
||||
uint32_t PreserveSRAFPRMask {};
|
||||
|
||||
@@ -1000,7 +983,7 @@ void Arm64Emitter::FillForPreserveAllABICall(bool FPRs) {
|
||||
PreserveSRAMask = x64::PreserveAll_SRAMask;
|
||||
PreserveSRAFPRMask = x64::PreserveAll_SRAFPRMask;
|
||||
|
||||
if (CanUseSVE) {
|
||||
if (CanUseSVE256) {
|
||||
DynamicFPRs = x64::PreserveAll_DynamicFPRSVE;
|
||||
PreserveSRAFPRMask = x64::PreserveAll_SRAFPRSVEMask;
|
||||
}
|
||||
@@ -1010,7 +993,7 @@ void Arm64Emitter::FillForPreserveAllABICall(bool FPRs) {
|
||||
PreserveSRAMask = x32::PreserveAll_SRAMask;
|
||||
PreserveSRAFPRMask = x32::PreserveAll_SRAFPRMask;
|
||||
|
||||
if (CanUseSVE) {
|
||||
if (CanUseSVE256) {
|
||||
DynamicFPRs = x32::PreserveAll_DynamicFPRSVE;
|
||||
PreserveSRAFPRMask = x32::PreserveAll_SRAFPRSVEMask;
|
||||
}
|
||||
@@ -1020,7 +1003,7 @@ void Arm64Emitter::FillForPreserveAllABICall(bool FPRs) {
|
||||
FillStaticRegs(true, PreserveSRAMask, PreserveSRAFPRMask);
|
||||
|
||||
// Pop the vector registers.
|
||||
PopVectorRegisters(CanUseSVE, DynamicFPRs);
|
||||
PopVectorRegisters(CanUseSVE256, DynamicFPRs);
|
||||
|
||||
// Pop the general registers.
|
||||
PopGeneralRegisters(DynamicGPRs);
|
||||
|
||||
@@ -2,9 +2,6 @@
|
||||
#pragma once
|
||||
|
||||
#include "FEXCore/Utils/EnumUtils.h"
|
||||
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
|
||||
#include "Interface/Core/ArchHelpers/CodeEmitter/Registers.h"
|
||||
|
||||
#include "Interface/Core/ObjectCache/Relocations.h"
|
||||
|
||||
#include <aarch64/assembler-aarch64.h>
|
||||
@@ -22,6 +19,8 @@
|
||||
|
||||
#include <FEXCore/Config/Config.h>
|
||||
#include <FEXCore/fextl/vector.h>
|
||||
#include <CodeEmitter/Emitter.h>
|
||||
#include <CodeEmitter/Registers.h>
|
||||
|
||||
#include <array>
|
||||
#include <cstddef>
|
||||
@@ -35,85 +34,74 @@ class ContextImpl;
|
||||
|
||||
namespace FEXCore::CPU {
|
||||
// Contains the address to the currently available CPU state
|
||||
constexpr auto STATE = FEXCore::ARMEmitter::XReg::x28;
|
||||
constexpr auto STATE = ARMEmitter::XReg::x28;
|
||||
|
||||
#ifndef _M_ARM_64EC
|
||||
// GPR temporaries. Only x3 can be used across spill boundaries
|
||||
// so if these ever need to change, be very careful about that.
|
||||
constexpr auto TMP1 = FEXCore::ARMEmitter::XReg::x0;
|
||||
constexpr auto TMP2 = FEXCore::ARMEmitter::XReg::x1;
|
||||
constexpr auto TMP3 = FEXCore::ARMEmitter::XReg::x2;
|
||||
constexpr auto TMP4 = FEXCore::ARMEmitter::XReg::x3;
|
||||
constexpr auto TMP1 = ARMEmitter::XReg::x0;
|
||||
constexpr auto TMP2 = ARMEmitter::XReg::x1;
|
||||
constexpr auto TMP3 = ARMEmitter::XReg::x2;
|
||||
constexpr auto TMP4 = ARMEmitter::XReg::x3;
|
||||
constexpr bool TMP_ABIARGS = true;
|
||||
|
||||
// We pin r26/r27 as PF/AF respectively, this is internal FEX ABI.
|
||||
constexpr auto REG_PF = FEXCore::ARMEmitter::Reg::r26;
|
||||
constexpr auto REG_AF = FEXCore::ARMEmitter::Reg::r27;
|
||||
constexpr auto REG_PF = ARMEmitter::Reg::r26;
|
||||
constexpr auto REG_AF = ARMEmitter::Reg::r27;
|
||||
|
||||
// Vector temporaries
|
||||
constexpr auto VTMP1 = FEXCore::ARMEmitter::VReg::v0;
|
||||
constexpr auto VTMP2 = FEXCore::ARMEmitter::VReg::v1;
|
||||
constexpr auto VTMP1 = ARMEmitter::VReg::v0;
|
||||
constexpr auto VTMP2 = ARMEmitter::VReg::v1;
|
||||
#else
|
||||
constexpr auto TMP1 = FEXCore::ARMEmitter::XReg::x10;
|
||||
constexpr auto TMP2 = FEXCore::ARMEmitter::XReg::x11;
|
||||
constexpr auto TMP3 = FEXCore::ARMEmitter::XReg::x12;
|
||||
constexpr auto TMP4 = FEXCore::ARMEmitter::XReg::x13;
|
||||
constexpr auto TMP1 = ARMEmitter::XReg::x10;
|
||||
constexpr auto TMP2 = ARMEmitter::XReg::x11;
|
||||
constexpr auto TMP3 = ARMEmitter::XReg::x12;
|
||||
constexpr auto TMP4 = ARMEmitter::XReg::x13;
|
||||
constexpr bool TMP_ABIARGS = false;
|
||||
|
||||
// We pin r11/r12 as PF/AF respectively for arm64ec, as r26/r27 are used for SRA.
|
||||
constexpr auto REG_PF = FEXCore::ARMEmitter::Reg::r9;
|
||||
constexpr auto REG_AF = FEXCore::ARMEmitter::Reg::r24;
|
||||
constexpr auto REG_PF = ARMEmitter::Reg::r9;
|
||||
constexpr auto REG_AF = ARMEmitter::Reg::r24;
|
||||
|
||||
// Vector temporaries
|
||||
constexpr auto VTMP1 = FEXCore::ARMEmitter::VReg::v16;
|
||||
constexpr auto VTMP2 = FEXCore::ARMEmitter::VReg::v17;
|
||||
constexpr auto VTMP1 = ARMEmitter::VReg::v16;
|
||||
constexpr auto VTMP2 = ARMEmitter::VReg::v17;
|
||||
|
||||
// Entry/Exit ABI
|
||||
constexpr auto EC_CALL_CHECKER_PC_REG = ARMEmitter::XReg::x9;
|
||||
constexpr auto EC_ENTRY_CPUAREA_REG = ARMEmitter::XReg::x17;
|
||||
#endif
|
||||
|
||||
// Predicate register temporaries (used when AVX support is enabled)
|
||||
// PRED_TMP_16B indicates a predicate register that indicates the first 16 bytes set to 1.
|
||||
// PRED_TMP_32B indicates a predicate register that indicates the first 32 bytes set to 1.
|
||||
constexpr FEXCore::ARMEmitter::PRegister PRED_TMP_16B = FEXCore::ARMEmitter::PReg::p6;
|
||||
constexpr FEXCore::ARMEmitter::PRegister PRED_TMP_32B = FEXCore::ARMEmitter::PReg::p7;
|
||||
constexpr ARMEmitter::PRegister PRED_TMP_16B = ARMEmitter::PReg::p6;
|
||||
constexpr ARMEmitter::PRegister PRED_TMP_32B = ARMEmitter::PReg::p7;
|
||||
|
||||
|
||||
// This class contains common emitter utility functions that can
|
||||
// be used by both Arm64 JIT and ARM64 Dispatcher
|
||||
class Arm64Emitter : public FEXCore::ARMEmitter::Emitter {
|
||||
class Arm64Emitter : public ARMEmitter::Emitter {
|
||||
protected:
|
||||
Arm64Emitter(FEXCore::Context::ContextImpl* ctx, void* EmissionPtr = nullptr, size_t size = 0);
|
||||
|
||||
FEXCore::Context::ContextImpl* EmitterCTX;
|
||||
vixl::aarch64::CPU CPU;
|
||||
|
||||
std::span<const FEXCore::ARMEmitter::Register> ConfiguredDynamicRegisterBase {};
|
||||
std::span<const FEXCore::ARMEmitter::Register> StaticRegisters {};
|
||||
std::span<const FEXCore::ARMEmitter::Register> GeneralRegisters {};
|
||||
std::span<const std::pair<FEXCore::ARMEmitter::Register, FEXCore::ARMEmitter::Register>> GeneralPairRegisters {};
|
||||
std::span<const FEXCore::ARMEmitter::VRegister> StaticFPRegisters {};
|
||||
std::span<const FEXCore::ARMEmitter::VRegister> GeneralFPRegisters {};
|
||||
std::span<const ARMEmitter::Register> ConfiguredDynamicRegisterBase {};
|
||||
std::span<const ARMEmitter::Register> StaticRegisters {};
|
||||
std::span<const ARMEmitter::Register> GeneralRegisters {};
|
||||
std::span<const ARMEmitter::VRegister> StaticFPRegisters {};
|
||||
std::span<const ARMEmitter::VRegister> GeneralFPRegisters {};
|
||||
uint32_t PairRegisters = 0;
|
||||
|
||||
/**
|
||||
* @name Register Allocation
|
||||
* @{ */
|
||||
constexpr static uint32_t RegisterClasses = 6;
|
||||
|
||||
constexpr static uint64_t GPRBase = (0ULL << 32);
|
||||
constexpr static uint64_t FPRBase = (1ULL << 32);
|
||||
constexpr static uint64_t GPRPairBase = (2ULL << 32);
|
||||
|
||||
/** @} */
|
||||
|
||||
constexpr static uint8_t RA_32 = 0;
|
||||
constexpr static uint8_t RA_64 = 1;
|
||||
constexpr static uint8_t RA_FPR = 2;
|
||||
|
||||
void LoadConstant(FEXCore::ARMEmitter::Size s, FEXCore::ARMEmitter::Register Reg, uint64_t Constant, bool NOPPad = false);
|
||||
void LoadConstant(ARMEmitter::Size s, ARMEmitter::Register Reg, uint64_t Constant, bool NOPPad = false);
|
||||
|
||||
|
||||
// NOTE: These functions WILL clobber the register TMP4 if AVX support is enabled
|
||||
// and FPRs are being spilled or filled. If only GPRs are spilled/filled, then
|
||||
// TMP4 is left alone.
|
||||
void SpillStaticRegs(FEXCore::ARMEmitter::Register TmpReg, bool FPRs = true, uint32_t GPRSpillMask = ~0U, uint32_t FPRSpillMask = ~0U);
|
||||
void SpillStaticRegs(ARMEmitter::Register TmpReg, bool FPRs = true, uint32_t GPRSpillMask = ~0U, uint32_t FPRSpillMask = ~0U);
|
||||
void FillStaticRegs(bool FPRs = true, uint32_t GPRFillMask = ~0U, uint32_t FPRFillMask = ~0U);
|
||||
|
||||
// Register 0-18 + 29 + 30 are caller saved
|
||||
@@ -124,13 +112,13 @@ protected:
|
||||
static constexpr uint32_t CALLER_FPR_MASK = ~0U;
|
||||
|
||||
// Generic push and pop vector registers.
|
||||
void PushVectorRegisters(FEXCore::ARMEmitter::Register TmpReg, bool SVERegs, std::span<const FEXCore::ARMEmitter::VRegister> VRegs);
|
||||
void PushGeneralRegisters(FEXCore::ARMEmitter::Register TmpReg, std::span<const FEXCore::ARMEmitter::Register> Regs);
|
||||
void PushVectorRegisters(ARMEmitter::Register TmpReg, bool SVERegs, std::span<const ARMEmitter::VRegister> VRegs);
|
||||
void PushGeneralRegisters(ARMEmitter::Register TmpReg, std::span<const ARMEmitter::Register> Regs);
|
||||
|
||||
void PopVectorRegisters(bool SVERegs, std::span<const FEXCore::ARMEmitter::VRegister> VRegs);
|
||||
void PopGeneralRegisters(std::span<const FEXCore::ARMEmitter::Register> Regs);
|
||||
void PopVectorRegisters(bool SVERegs, std::span<const ARMEmitter::VRegister> VRegs);
|
||||
void PopGeneralRegisters(std::span<const ARMEmitter::Register> Regs);
|
||||
|
||||
void PushDynamicRegsAndLR(FEXCore::ARMEmitter::Register TmpReg);
|
||||
void PushDynamicRegsAndLR(ARMEmitter::Register TmpReg);
|
||||
void PopDynamicRegsAndLR();
|
||||
|
||||
void PushCalleeSavedRegisters();
|
||||
@@ -146,10 +134,10 @@ protected:
|
||||
// Callee Saved:
|
||||
// - X9-X15, X19-X31
|
||||
// - Low 128-bits of v8-v31
|
||||
void SpillForPreserveAllABICall(FEXCore::ARMEmitter::Register TmpReg, bool FPRs = true);
|
||||
void SpillForPreserveAllABICall(ARMEmitter::Register TmpReg, bool FPRs = true);
|
||||
void FillForPreserveAllABICall(bool FPRs = true);
|
||||
|
||||
void SpillForABICall(bool SupportsPreserveAllABI, FEXCore::ARMEmitter::Register TmpReg, bool FPRs = true) {
|
||||
void SpillForABICall(bool SupportsPreserveAllABI, ARMEmitter::Register TmpReg, bool FPRs = true) {
|
||||
if (SupportsPreserveAllABI) {
|
||||
SpillForPreserveAllABICall(TmpReg, FPRs);
|
||||
} else {
|
||||
|
||||
File diff suppressed because it is too large.
Load diff
@@ -19,6 +19,10 @@ namespace CPU {
|
||||
{0x0000'0000'8000'0000ULL, 0x0000'0000'8000'0000ULL}, // NAMED_VECTOR_PADDSUBPS_INVERT_UPPER
|
||||
{0x8000'0000'0000'0000ULL, 0x0000'0000'0000'0000ULL}, // NAMED_VECTOR_PADDSUBPD_INVERT
|
||||
{0x8000'0000'0000'0000ULL, 0x0000'0000'0000'0000ULL}, // NAMED_VECTOR_PADDSUBPD_INVERT_UPPER
|
||||
{0x8000'0000'0000'0000ULL, 0x8000'0000'0000'0000ULL}, // NAMED_VECTOR_PSUBADDPS_INVERT
|
||||
{0x8000'0000'0000'0000ULL, 0x8000'0000'0000'0000ULL}, // NAMED_VECTOR_PSUBADDPS_INVERT_UPPER
|
||||
{0x0000'0000'0000'0000ULL, 0x8000'0000'0000'0000ULL}, // NAMED_VECTOR_PSUBADDPD_INVERT
|
||||
{0x0000'0000'0000'0000ULL, 0x8000'0000'0000'0000ULL}, // NAMED_VECTOR_PSUBADDPD_INVERT_UPPER
|
||||
{0x0000'0001'0000'0000ULL, 0x0000'0003'0000'0002ULL}, // NAMED_VECTOR_MOVMSKPS_SHIFT
|
||||
{0x040B'0E01'0B0E'0104ULL, 0x0C03'0609'0306'090CULL}, // NAMED_VECTOR_AESKEYGENASSIST_SWIZZLE
|
||||
{0x0706'0504'FFFF'FFFFULL, 0xFFFF'FFFF'0B0A'0908ULL}, // NAMED_VECTOR_BLENDPS_0110B
|
||||
|
||||
@@ -36,7 +36,6 @@ namespace CodeSerialize {
|
||||
namespace CPU {
|
||||
struct CPUBackendFeatures {
|
||||
bool SupportsFlags = false;
|
||||
bool SupportsSaturatingRoundingShifts = false;
|
||||
bool SupportsVTBL2 = false;
|
||||
};
|
||||
|
||||
@@ -140,7 +139,7 @@ namespace CPU {
|
||||
*/
|
||||
[[nodiscard]]
|
||||
virtual CompiledCode CompileCode(uint64_t Entry, const FEXCore::IR::IRListView* IR, FEXCore::Core::DebugData* DebugData,
|
||||
FEXCore::IR::RegisterAllocationData* RAData) = 0;
|
||||
const FEXCore::IR::RegisterAllocationData* RAData) = 0;
|
||||
|
||||
/**
|
||||
* @brief Relocates a block of code from the JIT code object cache
|
||||
|
||||
@@ -69,6 +69,8 @@ namespace ProductNames {
|
||||
|
||||
static const char ARM_Firestorm[] = "Apple Firestorm";
|
||||
static const char ARM_Icestorm[] = "Apple Icestorm";
|
||||
|
||||
static const char ARM_ORYON_1[] = "Oryon-1";
|
||||
#else
|
||||
#endif
|
||||
} // namespace ProductNames
|
||||
@@ -140,8 +142,10 @@ void CPUIDEmu::SetupHostHybridFlag() {
|
||||
// CPU priority order
|
||||
// This is mostly arbitrary but will sort by some sort of CPU priority by performance
|
||||
// Relative list so things they will commonly end up in big.little configurations sort of relate
|
||||
static constexpr std::array<CPUMIDR, 42> CPUMIDRs = {{
|
||||
static constexpr std::array<CPUMIDR, 43> CPUMIDRs = {{
|
||||
// Typically big CPU cores
|
||||
{0x51, 0x001, 1, ProductNames::ARM_ORYON_1}, // Qualcomm Oryon-1
|
||||
|
||||
{0x61, 0x023, 1, ProductNames::ARM_Firestorm}, // Apple M1 Firestorm
|
||||
|
||||
{0x41, 0xd82, 1, ProductNames::ARM_X4}, // X4
|
||||
@@ -343,8 +347,7 @@ void CPUIDEmu::SetupHostHybridFlag() {}
|
||||
|
||||
|
||||
void CPUIDEmu::SetupFeatures() {
|
||||
// TODO: Enable once AVX is supported.
|
||||
if (false && CTX->HostFeatures.SupportsAVX) {
|
||||
if (CTX->HostFeatures.SupportsAVX) {
|
||||
XCR0 |= XCR0_AVX;
|
||||
}
|
||||
|
||||
@@ -355,14 +358,14 @@ void CPUIDEmu::SetupFeatures() {
|
||||
return;
|
||||
}
|
||||
|
||||
#define ENABLE_DISABLE_OPTION(FeatureName, name, enum_name) \
|
||||
do { \
|
||||
const bool Disable##name = (CPUIDFeatures() & FEXCore::Config::CPUID::DISABLE##enum_name) != 0; \
|
||||
const bool Enable##name = (CPUIDFeatures() & FEXCore::Config::CPUID::ENABLE##enum_name) != 0; \
|
||||
#define ENABLE_DISABLE_OPTION(FeatureName, name, enum_name) \
|
||||
do { \
|
||||
const bool Disable##name = (CPUIDFeatures() & FEXCore::Config::CPUID::DISABLE##enum_name) != 0; \
|
||||
const bool Enable##name = (CPUIDFeatures() & FEXCore::Config::CPUID::ENABLE##enum_name) != 0; \
|
||||
LogMan::Throw::AFmt(!(Disable##name && Enable##name), "Disabling and Enabling CPU feature (" #name ") is mutually exclusive"); \
|
||||
const bool AlreadyEnabled = Features.FeatureName; \
|
||||
const bool Result = (AlreadyEnabled | Enable##name) & !Disable##name; \
|
||||
Features.FeatureName = Result; \
|
||||
const bool AlreadyEnabled = Features.FeatureName; \
|
||||
const bool Result = (AlreadyEnabled | Enable##name) & !Disable##name; \
|
||||
Features.FeatureName = Result; \
|
||||
} while (0)
|
||||
|
||||
ENABLE_DISABLE_OPTION(SHA, SHA, SHA);
|
||||
@@ -413,7 +416,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) const {
|
||||
(1 << 9) | // SSSE3
|
||||
(0 << 10) | // L1 context ID
|
||||
(0 << 11) | // Silicon debug
|
||||
(0 << 12) | // FMA3
|
||||
(SupportsAVX() << 12) | // FMA3
|
||||
(1 << 13) | // CMPXCHG16B
|
||||
(0 << 14) | // xTPR update control
|
||||
(0 << 15) | // Perfmon and debug capability
|
||||
@@ -430,7 +433,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_01h(uint32_t Leaf) const {
|
||||
(SupportsAVX() << 26) | // XSAVE
|
||||
(SupportsAVX() << 27) | // OSXSAVE
|
||||
(SupportsAVX() << 28) | // AVX
|
||||
(0 << 29) | // F16C
|
||||
(SupportsAVX() << 29) | // F16C
|
||||
(CTX->HostFeatures.SupportsRAND << 30) | // RDRAND
|
||||
(Hypervisor << 31);
|
||||
|
||||
@@ -597,6 +600,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) const {
|
||||
// This is due to LRCPC performance on Cortex being abysmal.
|
||||
// Only enable EnhancedREPMOVS if SoftwareTSO isn't required OR if MemcpySetTSO is not enabled.
|
||||
const uint32_t SupportsEnhancedREPMOVS = CTX->SoftwareTSORequired() == false || MemcpySetTSOEnabled() == false;
|
||||
const uint32_t SupportsVPCLMULQDQ = CTX->HostFeatures.SupportsPMULL_128Bit && SupportsAVX();
|
||||
|
||||
// Number of subfunctions
|
||||
Res.eax = 0x0;
|
||||
@@ -605,7 +609,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) const {
|
||||
(0 << 2) | // SGX
|
||||
(SupportsAVX() << 3) | // BMI1
|
||||
(0 << 4) | // Intel Hardware Lock Elison
|
||||
(0 << 5) | // AVX2 support
|
||||
(SupportsAVX() << 5) | // AVX2 support
|
||||
(1 << 6) | // FPU data pointer updated only on exception
|
||||
(1 << 7) | // SMEP support
|
||||
(SupportsAVX() << 8) | // BMI2
|
||||
@@ -633,38 +637,38 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_07h(uint32_t Leaf) const {
|
||||
(0 << 30) | // Reserved
|
||||
(0 << 31); // Reserved
|
||||
|
||||
Res.ecx = (1 << 0) | // PREFETCHWT1
|
||||
(0 << 1) | // AVX512VBMI
|
||||
(0 << 2) | // Usermode instruction prevention
|
||||
(0 << 3) | // Protection keys for user mode pages
|
||||
(0 << 4) | // OS protection keys
|
||||
(0 << 5) | // waitpkg
|
||||
(0 << 6) | // AVX512_VBMI2
|
||||
(0 << 7) | // CET shadow stack
|
||||
(0 << 8) | // GFNI
|
||||
(0 << 9) | // VAES
|
||||
(0 << 10) | // VPCLMULQDQ
|
||||
(0 << 11) | // AVX512_VNNI
|
||||
(0 << 12) | // AVX512_BITALG
|
||||
(0 << 13) | // Intel Total Memory Encryption
|
||||
(0 << 14) | // AVX512_VPOPCNTDQ
|
||||
(0 << 15) | // Reserved
|
||||
(0 << 16) | // 5 Level page tables
|
||||
(0 << 17) | // MPX MAWAU
|
||||
(0 << 18) | // MPX MAWAU
|
||||
(0 << 19) | // MPX MAWAU
|
||||
(0 << 20) | // MPX MAWAU
|
||||
(0 << 21) | // MPX MAWAU
|
||||
(1 << 22) | // RDPID Read Processor ID
|
||||
(0 << 23) | // Reserved
|
||||
(0 << 24) | // Reserved
|
||||
(0 << 25) | // CLDEMOTE
|
||||
(0 << 26) | // Reserved
|
||||
(0 << 27) | // MOVDIRI
|
||||
(0 << 28) | // MOVDIR64B
|
||||
(0 << 29) | // Reserved
|
||||
(0 << 30) | // SGX Launch configuration
|
||||
(0 << 31); // Reserved
|
||||
Res.ecx = (1 << 0) | // PREFETCHWT1
|
||||
(0 << 1) | // AVX512VBMI
|
||||
(0 << 2) | // Usermode instruction prevention
|
||||
(0 << 3) | // Protection keys for user mode pages
|
||||
(0 << 4) | // OS protection keys
|
||||
(0 << 5) | // waitpkg
|
||||
(0 << 6) | // AVX512_VBMI2
|
||||
(0 << 7) | // CET shadow stack
|
||||
(0 << 8) | // GFNI
|
||||
(CTX->HostFeatures.SupportsAES256 << 9) | // VAES
|
||||
(SupportsVPCLMULQDQ << 10) | // VPCLMULQDQ
|
||||
(0 << 11) | // AVX512_VNNI
|
||||
(0 << 12) | // AVX512_BITALG
|
||||
(0 << 13) | // Intel Total Memory Encryption
|
||||
(0 << 14) | // AVX512_VPOPCNTDQ
|
||||
(0 << 15) | // Reserved
|
||||
(0 << 16) | // 5 Level page tables
|
||||
(0 << 17) | // MPX MAWAU
|
||||
(0 << 18) | // MPX MAWAU
|
||||
(0 << 19) | // MPX MAWAU
|
||||
(0 << 20) | // MPX MAWAU
|
||||
(0 << 21) | // MPX MAWAU
|
||||
(1 << 22) | // RDPID Read Processor ID
|
||||
(0 << 23) | // Reserved
|
||||
(0 << 24) | // Reserved
|
||||
(0 << 25) | // CLDEMOTE
|
||||
(0 << 26) | // Reserved
|
||||
(0 << 27) | // MOVDIRI
|
||||
(0 << 28) | // MOVDIR64B
|
||||
(0 << 29) | // Reserved
|
||||
(0 << 30) | // SGX Launch configuration
|
||||
(0 << 31); // Reserved
|
||||
|
||||
Res.edx = (0 << 0) | // Reserved
|
||||
(0 << 1) | // Reserved
|
||||
@@ -883,7 +887,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0001h(uint32_t Leaf) con
|
||||
(0 << 18) | // Reserved
|
||||
(0 << 19) | // Reserved
|
||||
(0 << 20) | // Reserved
|
||||
(0 << 21) | // Reserved
|
||||
(0 << 21) | // XOP-TBM
|
||||
(0 << 22) | // Topology extensions support
|
||||
(0 << 23) | // Core performance counter extensions
|
||||
(0 << 24) | // NB performance counter extensions
|
||||
@@ -891,7 +895,7 @@ FEXCore::CPUID::FunctionResults CPUIDEmu::Function_8000_0001h(uint32_t Leaf) con
|
||||
(0 << 26) | // Data breakpoints extensions
|
||||
(0 << 27) | // Performance TSC
|
||||
(0 << 28) | // L2 perf counter extensions
|
||||
(0 << 29) | // Reserved
|
||||
(0 << 29) | // MONITORX
|
||||
(0 << 30) | // Reserved
|
||||
(0 << 31); // Reserved
|
||||
|
||||
|
||||
@@ -216,6 +216,55 @@ uint32_t ContextImpl::ReconstructCompactedEFLAGS(FEXCore::Core::InternalThreadSt
|
||||
return EFLAGS;
|
||||
}
|
||||
|
||||
void ContextImpl::ReconstructXMMRegisters(const FEXCore::Core::InternalThreadState* Thread, __uint128_t* XMM_Low, __uint128_t* YMM_High) {
|
||||
const size_t MaximumRegisters = Config.Is64BitMode ? FEXCore::Core::CPUState::NUM_XMMS : 8;
|
||||
|
||||
if (YMM_High != nullptr && HostFeatures.SupportsAVX) {
|
||||
const bool SupportsConvergedRegisters = HostFeatures.SupportsSVE256;
|
||||
|
||||
if (SupportsConvergedRegisters) {
|
||||
///< Output wants to de-interleave
|
||||
for (size_t i = 0; i < MaximumRegisters; ++i) {
|
||||
memcpy(&XMM_Low[i], &Thread->CurrentFrame->State.xmm.avx.data[i][0], sizeof(__uint128_t));
|
||||
memcpy(&YMM_High[i], &Thread->CurrentFrame->State.xmm.avx.data[i][2], sizeof(__uint128_t));
|
||||
}
|
||||
} else {
|
||||
///< Matches what FEX wants with non-converged registers
|
||||
for (size_t i = 0; i < MaximumRegisters; ++i) {
|
||||
memcpy(&XMM_Low[i], &Thread->CurrentFrame->State.xmm.sse.data[i][0], sizeof(__uint128_t));
|
||||
memcpy(&YMM_High[i], &Thread->CurrentFrame->State.avx_high[i][0], sizeof(__uint128_t));
|
||||
}
|
||||
}
|
||||
} else {
|
||||
// Only support SSE, no AVX here, even if requested.
|
||||
memcpy(XMM_Low, Thread->CurrentFrame->State.xmm.sse.data, MaximumRegisters * sizeof(__uint128_t));
|
||||
}
|
||||
}
|
||||
|
||||
void ContextImpl::SetXMMRegistersFromState(FEXCore::Core::InternalThreadState* Thread, const __uint128_t* XMM_Low, const __uint128_t* YMM_High) {
|
||||
const size_t MaximumRegisters = Config.Is64BitMode ? FEXCore::Core::CPUState::NUM_XMMS : 8;
|
||||
if (YMM_High != nullptr && HostFeatures.SupportsAVX) {
|
||||
const bool SupportsConvergedRegisters = HostFeatures.SupportsSVE256;
|
||||
|
||||
if (SupportsConvergedRegisters) {
|
||||
///< Output wants to de-interleave
|
||||
for (size_t i = 0; i < MaximumRegisters; ++i) {
|
||||
memcpy(&Thread->CurrentFrame->State.xmm.avx.data[i][0], &XMM_Low[i], sizeof(__uint128_t));
|
||||
memcpy(&Thread->CurrentFrame->State.xmm.avx.data[i][2], &YMM_High[i], sizeof(__uint128_t));
|
||||
}
|
||||
} else {
|
||||
///< Matches what FEX wants with non-converged registers
|
||||
for (size_t i = 0; i < MaximumRegisters; ++i) {
|
||||
memcpy(&Thread->CurrentFrame->State.xmm.sse.data[i][0], &XMM_Low[i], sizeof(__uint128_t));
|
||||
memcpy(&Thread->CurrentFrame->State.avx_high[i][0], &YMM_High[i], sizeof(__uint128_t));
|
||||
}
|
||||
}
|
||||
} else {
|
||||
// Only support SSE, no AVX here, even if requested.
|
||||
memcpy(Thread->CurrentFrame->State.xmm.sse.data, XMM_Low, MaximumRegisters * sizeof(__uint128_t));
|
||||
}
|
||||
}
|
||||
|
||||
void ContextImpl::SetFlagsFromCompactedEFLAGS(FEXCore::Core::InternalThreadState* Thread, uint32_t EFLAGS) {
|
||||
const auto Frame = Thread->CurrentFrame;
|
||||
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_EFLAG_BITS; ++i) {
|
||||
@@ -366,7 +415,7 @@ void ContextImpl::InitializeCompiler(FEXCore::Core::InternalThreadState* Thread)
|
||||
|
||||
Thread->CTX = this;
|
||||
|
||||
Thread->PassManager->AddDefaultPasses(this, Config.Core == FEXCore::Config::CONFIG_IRJIT);
|
||||
Thread->PassManager->AddDefaultPasses(this);
|
||||
Thread->PassManager->AddDefaultValidationPasses();
|
||||
|
||||
Thread->PassManager->RegisterSyscallHandler(SyscallHandler);
|
||||
@@ -374,7 +423,7 @@ void ContextImpl::InitializeCompiler(FEXCore::Core::InternalThreadState* Thread)
|
||||
// Create CPU backend
|
||||
switch (Config.Core) {
|
||||
case FEXCore::Config::CONFIG_IRJIT:
|
||||
Thread->PassManager->InsertRegisterAllocationPass(HostFeatures.SupportsAVX);
|
||||
Thread->PassManager->InsertRegisterAllocationPass();
|
||||
Thread->CPUBackend = FEXCore::CPU::CreateArm64JITCore(this, Thread);
|
||||
break;
|
||||
case FEXCore::Config::CONFIG_CUSTOM: Thread->CPUBackend = CustomCPUFactory(this, Thread); break;
|
||||
@@ -403,8 +452,6 @@ ContextImpl::CreateThread(uint64_t InitialRIP, uint64_t StackPointer, FEXCore::C
|
||||
InitializeCompiler(Thread);
|
||||
|
||||
Thread->CurrentFrame->State.DeferredSignalRefCount.Store(0);
|
||||
Thread->CurrentFrame->State.DeferredSignalFaultAddress =
|
||||
reinterpret_cast<Core::NonAtomicRefCounter<uint64_t>*>(FEXCore::Allocator::VirtualAlloc(4096));
|
||||
|
||||
if (Config.BlockJITNaming() || Config.GlobalJITNaming() || Config.LibraryJITNaming()) {
|
||||
// Allocate a JIT symbol buffer only if enabled.
|
||||
@@ -421,7 +468,8 @@ void ContextImpl::DestroyThread(FEXCore::Core::InternalThreadState* Thread, bool
|
||||
#endif
|
||||
}
|
||||
|
||||
FEXCore::Allocator::VirtualFree(reinterpret_cast<void*>(Thread->CurrentFrame->State.DeferredSignalFaultAddress), 4096);
|
||||
FEXCore::Allocator::VirtualProtect(&Thread->InterruptFaultPage, sizeof(Thread->InterruptFaultPage),
|
||||
Allocator::ProtectOptions::Read | Allocator::ProtectOptions::Write);
|
||||
delete Thread;
|
||||
}
|
||||
|
||||
@@ -469,29 +517,39 @@ static void IRDumper(FEXCore::Core::InternalThreadState* Thread, IR::IREmitter*
|
||||
fextl::fmt::print(FD, "IR-ShouldDump-{} 0x{:x}:\n{}\n@@@@@\n", RA ? "post" : "pre", GuestRIP, out.str());
|
||||
};
|
||||
|
||||
static void ValidateIR(ContextImpl* ctx, IR::IREmitter* IREmitter) {
|
||||
// Convert to text, Parse, Convert to text again and make sure the texts match
|
||||
fextl::stringstream out;
|
||||
static auto compaction = IR::CreateIRCompaction(ctx->OpDispatcherAllocator);
|
||||
compaction->Run(IREmitter);
|
||||
auto NewIR = IREmitter->ViewIR();
|
||||
Dump(&out, &NewIR, nullptr);
|
||||
out.seekg(0);
|
||||
FEXCore::Utils::PooledAllocatorMalloc Allocator;
|
||||
auto reparsed = IR::Parse(Allocator, out);
|
||||
if (reparsed == nullptr) {
|
||||
LOGMAN_MSG_A_FMT("Failed to parse IR\n");
|
||||
} else {
|
||||
fextl::stringstream out2;
|
||||
auto NewIR2 = reparsed->ViewIR();
|
||||
Dump(&out2, &NewIR2, nullptr);
|
||||
if (out.str() != out2.str()) {
|
||||
LogMan::Msg::IFmt("one:\n {}", out.str());
|
||||
LogMan::Msg::IFmt("two:\n {}", out2.str());
|
||||
LOGMAN_MSG_A_FMT("Parsed IR doesn't match\n");
|
||||
}
|
||||
// IRStorageBase with fully owned memory
|
||||
struct IRListCopy : public IR::IRStorageBase {
|
||||
std::span<std::byte> IRData;
|
||||
std::span<std::byte> ListData;
|
||||
|
||||
// TODO: Consider defaulting to empty RAData instead?
|
||||
IR::RegisterAllocationData::UniquePtr RADataInternal;
|
||||
|
||||
IRListCopy(const IR::IRListView& view, IR::RegisterAllocationData::UniquePtr RAData)
|
||||
: RADataInternal(std::move(RAData)) {
|
||||
std::byte* Storage = reinterpret_cast<std::byte*>(FEXCore::Allocator::malloc(view.GetDataSize() + view.GetListSize()));
|
||||
|
||||
IRData = {Storage, Storage + view.GetDataSize()};
|
||||
ListData = {Storage + view.GetDataSize(), Storage + view.GetDataSize() + view.GetListSize()};
|
||||
memcpy(IRData.data(), (char*)view.GetData(), IRData.size());
|
||||
memcpy(ListData.data(), (char*)view.GetListData(), ListData.size());
|
||||
}
|
||||
}
|
||||
|
||||
IRListCopy(const IRListCopy& other) = delete;
|
||||
IRListCopy(IRListCopy&& other) = delete;
|
||||
|
||||
~IRListCopy() {
|
||||
FEXCore::Allocator::free(IRData.data());
|
||||
}
|
||||
|
||||
const IR::RegisterAllocationData* RAData() override {
|
||||
return RADataInternal.get();
|
||||
}
|
||||
IR::IRListView GetIRView() override {
|
||||
return IR::IRListView {IRData.data(), ListData.data(), IRData.size(), ListData.size()};
|
||||
}
|
||||
};
|
||||
|
||||
|
||||
ContextImpl::GenerateIRResult
|
||||
ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, bool ExtendedDebugInfo, uint64_t MaxInst) {
|
||||
@@ -574,7 +632,7 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
|
||||
|
||||
Thread->OpDispatcher->SetCurrentCodeBlock(CodeWasChangedBlock);
|
||||
Thread->OpDispatcher->_ThreadRemoveCodeEntry();
|
||||
Thread->OpDispatcher->_ExitFunction(
|
||||
Thread->OpDispatcher->ExitFunction(
|
||||
Thread->OpDispatcher->_EntrypointOffset(IR::SizeToOpSize(GPRSize), Block.Entry + BlockInstructionsLength - GuestRIP));
|
||||
|
||||
auto NextOpBlock = Thread->OpDispatcher->CreateNewCodeBlockAfter(CurrentBlock);
|
||||
@@ -605,7 +663,7 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
|
||||
}
|
||||
// Invalid instruction
|
||||
Thread->OpDispatcher->InvalidOp(DecodedInfo);
|
||||
Thread->OpDispatcher->_ExitFunction(Thread->OpDispatcher->_EntrypointOffset(IR::SizeToOpSize(GPRSize), Block.Entry - GuestRIP));
|
||||
Thread->OpDispatcher->ExitFunction(Thread->OpDispatcher->_EntrypointOffset(IR::SizeToOpSize(GPRSize), Block.Entry - GuestRIP));
|
||||
}
|
||||
|
||||
const bool NeedsBlockEnd =
|
||||
@@ -615,14 +673,14 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
|
||||
if (HadDispatchError && TotalInstructions == 0) {
|
||||
// Couldn't handle any instruction in op dispatcher
|
||||
Thread->OpDispatcher->ResetWorkingList();
|
||||
return {nullptr, nullptr, 0, 0, 0, 0};
|
||||
return {nullptr, 0, 0, 0, 0};
|
||||
}
|
||||
|
||||
if (NeedsBlockEnd) {
|
||||
const uint8_t GPRSize = GetGPRSize();
|
||||
|
||||
// We had some instructions. Early exit
|
||||
Thread->OpDispatcher->_ExitFunction(
|
||||
Thread->OpDispatcher->ExitFunction(
|
||||
Thread->OpDispatcher->_EntrypointOffset(IR::SizeToOpSize(GPRSize), Block.Entry + BlockInstructionsLength - GuestRIP));
|
||||
break;
|
||||
}
|
||||
@@ -643,35 +701,26 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
|
||||
|
||||
auto ShouldDump = Thread->OpDispatcher->ShouldDumpIR();
|
||||
// Debug
|
||||
{
|
||||
if (ShouldDump) {
|
||||
IRDumper(Thread, IREmitter, GuestRIP, nullptr);
|
||||
}
|
||||
|
||||
if (static_cast<ContextImpl*>(Thread->CTX)->Config.ValidateIRarser) {
|
||||
ValidateIR(this, IREmitter);
|
||||
}
|
||||
if (ShouldDump) {
|
||||
IRDumper(Thread, IREmitter, GuestRIP, nullptr);
|
||||
}
|
||||
|
||||
// Run the passmanager over the IR from the dispatcher
|
||||
Thread->PassManager->Run(IREmitter);
|
||||
|
||||
// Debug
|
||||
{
|
||||
if (ShouldDump) {
|
||||
IRDumper(Thread, IREmitter, GuestRIP,
|
||||
Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->GetAllocationData() : nullptr);
|
||||
}
|
||||
if (ShouldDump) {
|
||||
IRDumper(Thread, IREmitter, GuestRIP,
|
||||
Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->GetAllocationData() : nullptr);
|
||||
}
|
||||
|
||||
auto RAData = Thread->PassManager->HasPass("RA") ? Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA")->PullAllocationData() : nullptr;
|
||||
auto IRList = IREmitter->CreateIRCopy();
|
||||
auto IRList = fextl::make_unique<IRListCopy>(IREmitter->ViewIR(), std::move(RAData));
|
||||
|
||||
IREmitter->DelayedDisownBuffer();
|
||||
|
||||
return {
|
||||
.IRList = IRList,
|
||||
.RAData = std::move(RAData),
|
||||
.IR = std::move(IRList),
|
||||
.TotalInstructions = TotalInstructions,
|
||||
.TotalInstructionsLength = TotalInstructionsLength,
|
||||
.StartAddr = Thread->FrontendDecoder->DecodedMinAddress,
|
||||
@@ -680,13 +729,6 @@ ContextImpl::GenerateIR(FEXCore::Core::InternalThreadState* Thread, uint64_t Gue
|
||||
}
|
||||
|
||||
ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, uint64_t MaxInst) {
|
||||
FEXCore::IR::IRListView* IRList {};
|
||||
FEXCore::Core::DebugData* DebugData {};
|
||||
FEXCore::IR::RegisterAllocationData::UniquePtr RAData {};
|
||||
bool GeneratedIR {};
|
||||
uint64_t StartAddr {};
|
||||
uint64_t Length {};
|
||||
|
||||
// JIT Code object cache lookup
|
||||
if (CodeObjectCacheService) {
|
||||
auto CodeCacheEntry = CodeObjectCacheService->FetchCodeObjectFromCache(GuestRIP);
|
||||
@@ -695,9 +737,8 @@ ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalT
|
||||
if (CompiledCode) {
|
||||
return {
|
||||
.CompiledCode = CompiledCode,
|
||||
.IRData = nullptr, // No IR data generated
|
||||
.IR = nullptr, // No IR/RA data generated
|
||||
.DebugData = nullptr, // nullptr here ensures that code serialization doesn't occur on from cache read
|
||||
.RAData = nullptr, // No RA data generated
|
||||
.GeneratedIR = false, // nullptr here ensures IR cache mechanisms won't run
|
||||
.StartAddr = 0, // Unused
|
||||
.Length = 0, // Unused
|
||||
@@ -713,49 +754,47 @@ ContextImpl::CompileCodeResult ContextImpl::CompileCode(FEXCore::Core::InternalT
|
||||
}
|
||||
}
|
||||
|
||||
fextl::unique_ptr<FEXCore::IR::IRStorageBase> IR;
|
||||
FEXCore::Core::DebugData* DebugData {};
|
||||
uint64_t StartAddr {};
|
||||
uint64_t Length {};
|
||||
|
||||
// AOT IR bookkeeping and cache
|
||||
{
|
||||
auto [IRCopy, RACopy, DebugDataCopy, _StartAddr, _Length, _GeneratedIR] = IRCaptureCache.PreGenerateIRFetch(Thread, GuestRIP, IRList);
|
||||
if (_GeneratedIR) {
|
||||
auto IRFromAOT = IRCaptureCache.PreGenerateIRFetch(Thread, GuestRIP);
|
||||
if (IRFromAOT) {
|
||||
// Setup pointers to internal structures
|
||||
IRList = IRCopy;
|
||||
RAData = std::move(RACopy);
|
||||
DebugData = DebugDataCopy;
|
||||
StartAddr = _StartAddr;
|
||||
Length = _Length;
|
||||
GeneratedIR = _GeneratedIR;
|
||||
IR = std::move(IRFromAOT->IR);
|
||||
DebugData = IRFromAOT->DebugData;
|
||||
StartAddr = IRFromAOT->StartAddr;
|
||||
Length = IRFromAOT->Length;
|
||||
}
|
||||
}
|
||||
|
||||
if (IRList == nullptr) {
|
||||
if (!IR) {
|
||||
// Generate IR + Meta Info
|
||||
auto [IRCopy, RACopy, TotalInstructions, TotalInstructionsLength, _StartAddr, _Length] =
|
||||
GenerateIR(Thread, GuestRIP, Config.GDBSymbols(), MaxInst);
|
||||
auto [IRCopy, TotalInstructions, TotalInstructionsLength, _StartAddr, _Length] = GenerateIR(Thread, GuestRIP, Config.GDBSymbols(), MaxInst);
|
||||
|
||||
// Setup pointers to internal structures
|
||||
IRList = IRCopy;
|
||||
RAData = std::move(RACopy);
|
||||
IR = std::move(IRCopy);
|
||||
DebugData = new FEXCore::Core::DebugData();
|
||||
StartAddr = _StartAddr;
|
||||
Length = _Length;
|
||||
|
||||
// These blocks aren't already in the cache
|
||||
GeneratedIR = true;
|
||||
}
|
||||
|
||||
if (IRList == nullptr) {
|
||||
if (!IR) {
|
||||
return {};
|
||||
}
|
||||
// Attempt to get the CPU backend to compile this code
|
||||
auto IRView = IR->GetIRView();
|
||||
return {
|
||||
// FEX currently throws away the CPUBackend::CompiledCode object other than the entrypoint
|
||||
// In the future with code caching getting wired up, we will pass the rest of the data forward.
|
||||
// TODO: Pass the data forward when code caching is wired up to this.
|
||||
.CompiledCode = Thread->CPUBackend->CompileCode(GuestRIP, IRList, DebugData, RAData.get()).BlockEntry,
|
||||
.IRData = IRList,
|
||||
.CompiledCode = Thread->CPUBackend->CompileCode(GuestRIP, &IRView, DebugData, IR->RAData()).BlockEntry,
|
||||
.IR = std::move(IR),
|
||||
.DebugData = DebugData,
|
||||
.RAData = std::move(RAData),
|
||||
.GeneratedIR = GeneratedIR,
|
||||
.GeneratedIR = true,
|
||||
.StartAddr = StartAddr,
|
||||
.Length = Length,
|
||||
};
|
||||
@@ -782,21 +821,7 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
|
||||
return HostCode;
|
||||
}
|
||||
|
||||
void* CodePtr {};
|
||||
FEXCore::IR::IRListView* IRList {};
|
||||
FEXCore::Core::DebugData* DebugData {};
|
||||
|
||||
bool GeneratedIR {};
|
||||
uint64_t StartAddr {}, Length {};
|
||||
|
||||
auto [Code, IR, Data, RAData, Generated, _StartAddr, _Length] = CompileCode(Thread, GuestRIP, MaxInst);
|
||||
CodePtr = Code;
|
||||
IRList = IR;
|
||||
DebugData = Data;
|
||||
GeneratedIR = Generated;
|
||||
StartAddr = _StartAddr;
|
||||
Length = _Length;
|
||||
|
||||
auto [CodePtr, IR, DebugData, GeneratedIR, StartAddr, Length] = CompileCode(Thread, GuestRIP, MaxInst);
|
||||
if (CodePtr == nullptr) {
|
||||
return 0;
|
||||
}
|
||||
@@ -847,7 +872,7 @@ uintptr_t ContextImpl::CompileBlock(FEXCore::Core::CpuStateFrame* Frame, uint64_
|
||||
// Clear any relocations that might have been generated
|
||||
Thread->CPUBackend->ClearRelocations();
|
||||
|
||||
if (IRCaptureCache.PostCompileCode(Thread, CodePtr, GuestRIP, StartAddr, Length, std::move(RAData), IRList, DebugData, GeneratedIR)) {
|
||||
if (IRCaptureCache.PostCompileCode(Thread, CodePtr, GuestRIP, StartAddr, Length, std::move(IR), DebugData, GeneratedIR)) {
|
||||
// Early exit
|
||||
return (uintptr_t)CodePtr;
|
||||
}
|
||||
|
||||
@@ -1,7 +1,6 @@
|
||||
// SPDX-License-Identifier: MIT
|
||||
|
||||
#include "Interface/Context/Context.h"
|
||||
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
|
||||
#include "Interface/Core/Dispatcher/Dispatcher.h"
|
||||
#include "Interface/Core/LookupCache.h"
|
||||
#include "Interface/Core/X86HelperGen.h"
|
||||
@@ -17,6 +16,8 @@
|
||||
#include <FEXCore/Utils/LogManager.h>
|
||||
#include <FEXCore/Utils/MathUtils.h>
|
||||
|
||||
#include <CodeEmitter/Emitter.h>
|
||||
|
||||
#include <atomic>
|
||||
#include <condition_variable>
|
||||
#include <csignal>
|
||||
@@ -62,6 +63,11 @@ void Dispatcher::EmitDispatcher() {
|
||||
ARMEmitter::ForwardLabel l_CTX;
|
||||
ARMEmitter::SingleUseForwardLabel l_Sleep;
|
||||
#ifdef _M_ARM_64EC
|
||||
// These structures are not included in the standard Windows headers, define them here
|
||||
static constexpr size_t TEBCPUAreaOffset = 0x1788;
|
||||
static constexpr size_t CPUAreaInSyscallCallbackOffset = 0x1;
|
||||
static constexpr size_t CPUAreaEmulatorStackLimitOffset = 0x8;
|
||||
static constexpr size_t CPUAreaEmulatorDataOffset = 0x30;
|
||||
ARMEmitter::SingleUseForwardLabel ExitEC;
|
||||
#endif
|
||||
ARMEmitter::SingleUseForwardLabel l_CompileBlock;
|
||||
@@ -82,20 +88,45 @@ void Dispatcher::EmitDispatcher() {
|
||||
AbsoluteLoopTopAddressFillSRA = GetCursorAddress<uint64_t>();
|
||||
|
||||
FillStaticRegs();
|
||||
ARMEmitter::BiDirectionalLabel LoopTop {};
|
||||
|
||||
#ifdef _M_ARM_64EC
|
||||
b(&LoopTop);
|
||||
|
||||
AbsoluteLoopTopAddressEnterECFillSRA = GetCursorAddress<uint64_t>();
|
||||
ldr(STATE, EC_ENTRY_CPUAREA_REG, CPUAreaEmulatorDataOffset);
|
||||
FillStaticRegs();
|
||||
|
||||
// Enter JIT
|
||||
b(&LoopTop);
|
||||
|
||||
AbsoluteLoopTopAddressEnterEC = GetCursorAddress<uint64_t>();
|
||||
// Load ThreadState and write the target PC there
|
||||
ldr(STATE, EC_ENTRY_CPUAREA_REG, CPUAreaEmulatorDataOffset);
|
||||
str(EC_CALL_CHECKER_PC_REG, STATE_PTR(CpuStateFrame, State.rip));
|
||||
|
||||
// Swap stacks to the emulator stack
|
||||
ldr(TMP1, EC_ENTRY_CPUAREA_REG, CPUAreaEmulatorStackLimitOffset);
|
||||
add(ARMEmitter::Size::i64Bit, StaticRegisters[X86State::REG_RSP], ARMEmitter::Reg::rsp, 0);
|
||||
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, TMP1, 0);
|
||||
|
||||
if (EmitterCTX->HostFeatures.SupportsSVE128) {
|
||||
ptrue(ARMEmitter::SubRegSize::i8Bit, PRED_TMP_16B, ARMEmitter::PredicatePattern::SVE_VL16);
|
||||
}
|
||||
|
||||
// Enter JIT
|
||||
#endif
|
||||
|
||||
// We want to ensure that we are 16 byte aligned at the top of this loop
|
||||
Align16B();
|
||||
ARMEmitter::BiDirectionalLabel FullLookup {};
|
||||
ARMEmitter::BiDirectionalLabel CallBlock {};
|
||||
ARMEmitter::BackwardLabel LoopTop {};
|
||||
|
||||
Bind(&LoopTop);
|
||||
AbsoluteLoopTopAddress = GetCursorAddress<uint64_t>();
|
||||
|
||||
// Load in our RIP
|
||||
// Don't modify TMP3 since it contains our RIP once the block doesn't exist
|
||||
// IMPORTANT: Pointers.Common.ExitFunctionEC callsites/implementations need to be
|
||||
// adjusted accordingly if this changes.
|
||||
auto RipReg = TMP3;
|
||||
ldr(RipReg, STATE_PTR(CpuStateFrame, State.rip));
|
||||
|
||||
@@ -177,7 +208,8 @@ void Dispatcher::EmitDispatcher() {
|
||||
#ifdef _M_ARM_64EC
|
||||
{
|
||||
Bind(&ExitEC);
|
||||
// Target PC is already loaded into TMP3 at the start of the dispatcher
|
||||
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, StaticRegisters[X86State::REG_RSP], 0);
|
||||
mov(EC_CALL_CHECKER_PC_REG, RipReg);
|
||||
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionEC));
|
||||
br(TMP2);
|
||||
}
|
||||
@@ -204,6 +236,12 @@ void Dispatcher::EmitDispatcher() {
|
||||
add(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::x0, ARMEmitter::XReg::x0, 1);
|
||||
str(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
|
||||
|
||||
#ifdef _M_ARM_64EC
|
||||
ldr(ARMEmitter::XReg::x0, ARMEmitter::XReg::x18, TEBCPUAreaOffset);
|
||||
LoadConstant(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r1, 1);
|
||||
strb(ARMEmitter::WReg::w1, ARMEmitter::XReg::x0, CPUAreaInSyscallCallbackOffset);
|
||||
#endif
|
||||
|
||||
mov(ARMEmitter::XReg::x0, STATE);
|
||||
mov(ARMEmitter::XReg::x1, ARMEmitter::XReg::lr);
|
||||
|
||||
@@ -220,13 +258,18 @@ void Dispatcher::EmitDispatcher() {
|
||||
|
||||
FillStaticRegs();
|
||||
|
||||
#ifdef _M_ARM_64EC
|
||||
ldr(TMP2, ARMEmitter::XReg::x18, TEBCPUAreaOffset);
|
||||
strb(ARMEmitter::WReg::zr, TMP2, CPUAreaInSyscallCallbackOffset);
|
||||
#endif
|
||||
|
||||
ldr(TMP2, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
|
||||
sub(ARMEmitter::Size::i64Bit, TMP2, TMP2, 1);
|
||||
str(TMP2, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
|
||||
|
||||
// Trigger segfault if any deferred signals are pending
|
||||
ldr(TMP2, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalFaultAddress));
|
||||
str(ARMEmitter::XReg::zr, TMP2, 0);
|
||||
strb(ARMEmitter::XReg::zr, STATE,
|
||||
offsetof(FEXCore::Core::InternalThreadState, InterruptFaultPage) - offsetof(FEXCore::Core::InternalThreadState, BaseFrameState));
|
||||
|
||||
br(TMP1);
|
||||
}
|
||||
@@ -245,6 +288,12 @@ void Dispatcher::EmitDispatcher() {
|
||||
add(ARMEmitter::Size::i64Bit, ARMEmitter::XReg::x0, ARMEmitter::XReg::x0, 1);
|
||||
str(ARMEmitter::XReg::x0, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
|
||||
|
||||
#ifdef _M_ARM_64EC
|
||||
ldr(ARMEmitter::XReg::x0, ARMEmitter::XReg::x18, TEBCPUAreaOffset);
|
||||
LoadConstant(ARMEmitter::Size::i32Bit, ARMEmitter::Reg::r1, 1);
|
||||
strb(ARMEmitter::WReg::w1, ARMEmitter::XReg::x0, CPUAreaInSyscallCallbackOffset);
|
||||
#endif
|
||||
|
||||
ldr(ARMEmitter::XReg::x0, &l_CTX);
|
||||
mov(ARMEmitter::XReg::x1, STATE);
|
||||
// x2 contains guest RIP
|
||||
@@ -259,13 +308,18 @@ void Dispatcher::EmitDispatcher() {
|
||||
|
||||
FillStaticRegs();
|
||||
|
||||
#ifdef _M_ARM_64EC
|
||||
ldr(TMP1, ARMEmitter::XReg::x18, TEBCPUAreaOffset);
|
||||
strb(ARMEmitter::WReg::zr, TMP1, CPUAreaInSyscallCallbackOffset);
|
||||
#endif
|
||||
|
||||
ldr(TMP1, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
|
||||
sub(ARMEmitter::Size::i64Bit, TMP1, TMP1, 1);
|
||||
str(TMP1, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount));
|
||||
|
||||
// Trigger segfault if any deferred signals are pending
|
||||
ldr(TMP1, STATE, offsetof(FEXCore::Core::CPUState, DeferredSignalFaultAddress));
|
||||
str(ARMEmitter::XReg::zr, TMP1, 0);
|
||||
strb(ARMEmitter::XReg::zr, STATE,
|
||||
offsetof(FEXCore::Core::InternalThreadState, InterruptFaultPage) - offsetof(FEXCore::Core::InternalThreadState, BaseFrameState));
|
||||
|
||||
b(&LoopTop);
|
||||
}
|
||||
@@ -494,6 +548,8 @@ void Dispatcher::InitThreadPointers(FEXCore::Core::InternalThreadState* Thread)
|
||||
|
||||
Common.DispatcherLoopTop = AbsoluteLoopTopAddress;
|
||||
Common.DispatcherLoopTopFillSRA = AbsoluteLoopTopAddressFillSRA;
|
||||
Common.DispatcherLoopTopEnterEC = AbsoluteLoopTopAddressEnterEC;
|
||||
Common.DispatcherLoopTopEnterECFillSRA = AbsoluteLoopTopAddressEnterECFillSRA;
|
||||
Common.ExitFunctionLinker = ExitFunctionLinkerAddress;
|
||||
Common.ThreadStopHandlerSpillSRA = ThreadStopHandlerAddressSpillSRA;
|
||||
Common.ThreadPauseHandlerSpillSRA = ThreadPauseHandlerAddressSpillSRA;
|
||||
|
||||
@@ -47,6 +47,8 @@ public:
|
||||
uint64_t ThreadStopHandlerAddressSpillSRA {};
|
||||
uint64_t AbsoluteLoopTopAddress {};
|
||||
uint64_t AbsoluteLoopTopAddressFillSRA {};
|
||||
uint64_t AbsoluteLoopTopAddressEnterEC {};
|
||||
uint64_t AbsoluteLoopTopAddressEnterECFillSRA {};
|
||||
uint64_t ThreadPauseHandlerAddress {};
|
||||
uint64_t ThreadPauseHandlerAddressSpillSRA {};
|
||||
uint64_t ExitFunctionLinkerAddress {};
|
||||
|
||||
@@ -221,14 +221,16 @@ void Decoder::DecodeModRM_64(X86Tables::DecodedOperand* Operand, X86Tables::ModR
|
||||
{
|
||||
// If we have a VSIB byte (as opposed to SIB), then the index register is a vector.
|
||||
const bool IsIndexVector = (DecodeInst->TableInfo->Flags & InstFlags::FLAGS_VEX_VSIB) != 0;
|
||||
uint8_t InvalidSIBIndex = 0b100; ///< SIB Index where there is no register encoding.
|
||||
if (IsIndexVector) {
|
||||
DecodeInst->Flags |= X86Tables::DecodeFlags::FLAG_VSIB_BYTE;
|
||||
InvalidSIBIndex = ~0; ///< No Invalid SIB Index with Index Vectors.
|
||||
}
|
||||
|
||||
const uint8_t IndexREX = (DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_X) != 0 ? 1 : 0;
|
||||
const uint8_t BaseREX = (DecodeInst->Flags & DecodeFlags::FLAG_REX_XGPR_B) != 0 ? 1 : 0;
|
||||
|
||||
Operand->Data.SIB.Index = MapModRMToReg(IndexREX, SIB.index, false, false, IsIndexVector, false, 0b100);
|
||||
Operand->Data.SIB.Index = MapModRMToReg(IndexREX, SIB.index, false, false, IsIndexVector, false, InvalidSIBIndex);
|
||||
Operand->Data.SIB.Base = MapModRMToReg(BaseREX, SIB.base, false, false, false, false, ModRM.mod == 0 ? 0b101 : 16);
|
||||
}
|
||||
|
||||
@@ -659,6 +661,9 @@ bool Decoder::NormalOpHeader(const FEXCore::X86Tables::X86InstInfo* Info, uint16
|
||||
if (CTX->Config.Is64BitMode && (Byte1 & 0b00100000) == 0) {
|
||||
DecodeInst->Flags |= DecodeFlags::FLAG_REX_XGPR_B;
|
||||
}
|
||||
if (options.w) {
|
||||
DecodeInst->Flags |= DecodeFlags::FLAG_OPTION_AVX_W;
|
||||
}
|
||||
if (!(map_select >= 1 && map_select <= 3)) {
|
||||
LogMan::Msg::EFmt("We don't understand a map_select of: {}", map_select);
|
||||
return false;
|
||||
@@ -941,14 +946,12 @@ void Decoder::BranchTargetInMultiblockRange() {
|
||||
// auto RIPOffset = LoadSource(Op, Op->Src[0], Op->Flags);
|
||||
// auto RIPTargetConst = _Constant(Op->PC + Op->InstSize);
|
||||
// Target offset is PC + InstSize + Literal
|
||||
LOGMAN_THROW_A_FMT(DecodeInst->Src[0].IsLiteral(), "Had wrong operand type");
|
||||
TargetRIP = DecodeInst->PC + DecodeInst->InstSize + DecodeInst->Src[0].Data.Literal.Value;
|
||||
TargetRIP = DecodeInst->PC + DecodeInst->InstSize + DecodeInst->Src[0].Literal();
|
||||
break;
|
||||
}
|
||||
case 0xE9:
|
||||
case 0xEB: // Both are unconditional JMP instructions
|
||||
LOGMAN_THROW_A_FMT(DecodeInst->Src[0].IsLiteral(), "Had wrong operand type");
|
||||
TargetRIP = DecodeInst->PC + DecodeInst->InstSize + DecodeInst->Src[0].Data.Literal.Value;
|
||||
TargetRIP = DecodeInst->PC + DecodeInst->InstSize + DecodeInst->Src[0].Literal();
|
||||
Conditional = false;
|
||||
break;
|
||||
case 0xE8: // Call - Immediate target, We don't want to inline calls
|
||||
@@ -1000,8 +1003,7 @@ bool Decoder::BranchTargetCanContinue(bool FinalInstruction) const {
|
||||
|
||||
if (DecodeInst->OP == 0xE8) { // Call - immediate target
|
||||
const uint64_t NextRIP = DecodeInst->PC + DecodeInst->InstSize;
|
||||
LOGMAN_THROW_A_FMT(DecodeInst->Src[0].IsLiteral(), "Had wrong operand type");
|
||||
TargetRIP = DecodeInst->PC + DecodeInst->InstSize + DecodeInst->Src[0].Data.Literal.Value;
|
||||
TargetRIP = DecodeInst->PC + DecodeInst->InstSize + DecodeInst->Src[0].Literal();
|
||||
|
||||
if (GPRSize == 4) {
|
||||
// If we are running a 32bit guest then wrap around addresses that go above 32bit
|
||||
|
||||
@@ -44,6 +44,13 @@ static uint32_t GetFPCR() {
|
||||
static void SetFPCR(uint64_t Value) {
|
||||
__asm("msr FPCR, %[Value]" ::[Value] "r"(Value));
|
||||
}
|
||||
|
||||
static uint32_t GetMIDR() {
|
||||
uint64_t Result {};
|
||||
__asm("mrs %[Res], MIDR_EL1" : [Res] "=r"(Result));
|
||||
return Result;
|
||||
}
|
||||
|
||||
#else
|
||||
static uint32_t GetDCZID() {
|
||||
// Return unsupported
|
||||
@@ -51,7 +58,7 @@ static uint32_t GetDCZID() {
|
||||
}
|
||||
#endif
|
||||
|
||||
static void OverrideFeatures(HostFeatures* Features) {
|
||||
static void OverrideFeatures(HostFeatures* Features, uint64_t ForceSVEWidth) {
|
||||
// Override features if the user has specifically called for it.
|
||||
FEX_CONFIG_OPT(HostFeatures, HOSTFEATURES);
|
||||
if (!HostFeatures()) {
|
||||
@@ -59,24 +66,23 @@ static void OverrideFeatures(HostFeatures* Features) {
|
||||
return;
|
||||
}
|
||||
|
||||
#define ENABLE_DISABLE_OPTION(FeatureName, name, enum_name) \
|
||||
do { \
|
||||
const bool Disable##name = (HostFeatures() & FEXCore::Config::HostFeatures::DISABLE##enum_name) != 0; \
|
||||
const bool Enable##name = (HostFeatures() & FEXCore::Config::HostFeatures::ENABLE##enum_name) != 0; \
|
||||
#define ENABLE_DISABLE_OPTION(FeatureName, name, enum_name) \
|
||||
do { \
|
||||
const bool Disable##name = (HostFeatures() & FEXCore::Config::HostFeatures::DISABLE##enum_name) != 0; \
|
||||
const bool Enable##name = (HostFeatures() & FEXCore::Config::HostFeatures::ENABLE##enum_name) != 0; \
|
||||
LogMan::Throw::AFmt(!(Disable##name && Enable##name), "Disabling and Enabling CPU feature (" #name ") is mutually exclusive"); \
|
||||
const bool AlreadyEnabled = Features->FeatureName; \
|
||||
const bool Result = (AlreadyEnabled | Enable##name) & !Disable##name; \
|
||||
Features->FeatureName = Result; \
|
||||
const bool AlreadyEnabled = Features->FeatureName; \
|
||||
const bool Result = (AlreadyEnabled | Enable##name) & !Disable##name; \
|
||||
Features->FeatureName = Result; \
|
||||
} while (0)
|
||||
|
||||
#define GET_SINGLE_OPTION(name, enum_name) \
|
||||
#define GET_SINGLE_OPTION(name, enum_name) \
|
||||
const bool Disable##name = (HostFeatures() & FEXCore::Config::HostFeatures::DISABLE##enum_name) != 0; \
|
||||
const bool Enable##name = (HostFeatures() & FEXCore::Config::HostFeatures::ENABLE##enum_name) != 0; \
|
||||
const bool Enable##name = (HostFeatures() & FEXCore::Config::HostFeatures::ENABLE##enum_name) != 0; \
|
||||
LogMan::Throw::AFmt(!(Disable##name && Enable##name), "Disabling and Enabling CPU feature (" #name ") is mutually exclusive");
|
||||
|
||||
ENABLE_DISABLE_OPTION(SupportsAVX, AVX, AVX);
|
||||
ENABLE_DISABLE_OPTION(SupportsAVX2, AVX2, AVX2);
|
||||
ENABLE_DISABLE_OPTION(SupportsSVE, SVE, SVE);
|
||||
ENABLE_DISABLE_OPTION(SupportsSVE128, SVE, SVE);
|
||||
ENABLE_DISABLE_OPTION(SupportsAFP, AFP, AFP);
|
||||
ENABLE_DISABLE_OPTION(SupportsRCPC, LRCPC, LRCPC);
|
||||
ENABLE_DISABLE_OPTION(SupportsTSOImm9, LRCPC2, LRCPC2);
|
||||
@@ -100,12 +106,17 @@ static void OverrideFeatures(HostFeatures* Features) {
|
||||
Features->SupportsCRC = true;
|
||||
Features->SupportsSHA = true;
|
||||
Features->SupportsPMULL_128Bit = true;
|
||||
Features->SupportsAES256 = true;
|
||||
} else if (DisableCrypto) {
|
||||
Features->SupportsAES = false;
|
||||
Features->SupportsCRC = false;
|
||||
Features->SupportsSHA = false;
|
||||
Features->SupportsPMULL_128Bit = false;
|
||||
Features->SupportsAES256 = false;
|
||||
}
|
||||
|
||||
///< Only force enable SVE256 if SVE is already enabled and ForceSVEWidth is set to >= 256.
|
||||
Features->SupportsSVE256 = ForceSVEWidth && ForceSVEWidth >= 256;
|
||||
}
|
||||
|
||||
HostFeatures::HostFeatures() {
|
||||
@@ -122,6 +133,9 @@ HostFeatures::HostFeatures() {
|
||||
auto Features = vixl::CPUFeatures::InferFromIDRegisters();
|
||||
#endif
|
||||
|
||||
FEX_CONFIG_OPT(ForceSVEWidth, FORCESVEWIDTH);
|
||||
FEX_CONFIG_OPT(Is64BitMode, IS64BIT_MODE);
|
||||
|
||||
SupportsAES = Features.Has(vixl::CPUFeatures::Feature::kAES);
|
||||
SupportsCRC = Features.Has(vixl::CPUFeatures::Feature::kCRC32);
|
||||
SupportsSHA = Features.Has(vixl::CPUFeatures::Feature::kSHA1) && Features.Has(vixl::CPUFeatures::Feature::kSHA2);
|
||||
@@ -141,25 +155,23 @@ HostFeatures::HostFeatures() {
|
||||
|
||||
Supports3DNow = true;
|
||||
SupportsSSE4A = true;
|
||||
|
||||
#ifdef VIXL_SIMULATOR
|
||||
// Hardcode enable SVE with 256-bit wide registers.
|
||||
SupportsSVE = true;
|
||||
SupportsAVX = true;
|
||||
SupportsSVE128 = ForceSVEWidth() ? ForceSVEWidth() >= 128 : true;
|
||||
SupportsSVE256 = ForceSVEWidth() ? ForceSVEWidth() >= 256 : true;
|
||||
#else
|
||||
SupportsSVE = Features.Has(vixl::CPUFeatures::Feature::kSVE);
|
||||
SupportsAVX = Features.Has(vixl::CPUFeatures::Feature::kSVE2) && vixl::aarch64::CPU::ReadSVEVectorLengthInBits() >= 256;
|
||||
SupportsSVE128 = Features.Has(vixl::CPUFeatures::Feature::kSVE2);
|
||||
SupportsSVE256 = Features.Has(vixl::CPUFeatures::Feature::kSVE2) && vixl::aarch64::CPU::ReadSVEVectorLengthInBits() >= 256;
|
||||
#endif
|
||||
// TODO: AVX2 is currently unsupported. Disable until the remaining features are implemented.
|
||||
SupportsAVX2 = false;
|
||||
SupportsAVX = true;
|
||||
|
||||
SupportsAES256 = SupportsAVX && SupportsAES;
|
||||
|
||||
SupportsBMI1 = true;
|
||||
SupportsBMI2 = true;
|
||||
SupportsCLWB = true;
|
||||
|
||||
// TODO: AFP is disabled until the scalar usage in the codebase can be audited to be working as expected.
|
||||
SupportsAFP = false;
|
||||
// RPRES has a dependency on AFP. Disable it until AFP is enabled.
|
||||
SupportsRPRES = false;
|
||||
|
||||
if (!SupportsAtomics) {
|
||||
WARN_ONCE_FMT("Host CPU doesn't support atomics. Expect bad performance");
|
||||
}
|
||||
@@ -190,6 +202,24 @@ HostFeatures::HostFeatures() {
|
||||
|
||||
// Set FPCR back to original just in case anything changed
|
||||
SetFPCR(OriginalFPCR);
|
||||
|
||||
if (SupportsRAND) {
|
||||
const auto MIDR = GetMIDR();
|
||||
constexpr uint32_t Implementer_QCOM = 0x51;
|
||||
constexpr uint32_t PartNum_Oryon1 = 0x001;
|
||||
const uint32_t MIDR_Implementer = (MIDR >> 24) & 0xFF;
|
||||
const uint32_t MIDR_PartNum = (MIDR >> 4) & 0xFFF;
|
||||
if (MIDR_Implementer == Implementer_QCOM && MIDR_PartNum == PartNum_Oryon1) {
|
||||
// Work around an errata in Qualcomm's Oryon.
|
||||
// While this CPU implements the RAND extension:
|
||||
// - The RNDR register works.
|
||||
// - The RNDRRS register will never read a random number. (Always return failure)
|
||||
// This is contrary to x86 RNG behaviour where it allows spurious failure with RDSEED, but guarantees eventual success.
|
||||
// This manifested itself on Linux when an x86 processor failed to guarantee forward progress and boot of services would infinite
|
||||
// loop. Just disable this extension if this CPU is detected.
|
||||
SupportsRAND = false;
|
||||
}
|
||||
}
|
||||
#endif
|
||||
|
||||
#ifdef VIXL_SIMULATOR
|
||||
@@ -224,12 +254,12 @@ HostFeatures::HostFeatures() {
|
||||
Supports3DNow = X86Features.has(Xbyak::util::Cpu::t3DN) && X86Features.has(Xbyak::util::Cpu::tE3DN);
|
||||
SupportsSSE4A = X86Features.has(Xbyak::util::Cpu::tSSE4a);
|
||||
SupportsAVX = true;
|
||||
SupportsAVX2 = true;
|
||||
SupportsSHA = X86Features.has(Xbyak::util::Cpu::tSHA);
|
||||
SupportsBMI1 = X86Features.has(Xbyak::util::Cpu::tBMI1);
|
||||
SupportsBMI2 = X86Features.has(Xbyak::util::Cpu::tBMI2);
|
||||
SupportsCLWB = X86Features.has(Xbyak::util::Cpu::tCLWB);
|
||||
SupportsPMULL_128Bit = X86Features.has(Xbyak::util::Cpu::tPCLMULQDQ);
|
||||
SupportsAES256 = SupportsAES && X86Features.has(Xbyak::util::Cpu::tVAES);
|
||||
|
||||
// xbyak doesn't know how to check for CLZero
|
||||
// First ensure we support a new enough extended CPUID function range
|
||||
@@ -247,6 +277,17 @@ HostFeatures::HostFeatures() {
|
||||
#endif
|
||||
#endif
|
||||
SupportsPreserveAllABI = FEXCORE_HAS_PRESERVE_ALL_ATTR;
|
||||
OverrideFeatures(this);
|
||||
|
||||
if (!Is64BitMode()) {
|
||||
///< Always disable AVX and AVX2 in 32-bit mode.
|
||||
// When AVX256 is enabled, signal frames start using significantly more stack space.
|
||||
// - 16bytes * 16 registers = 256 bytes for XMM registers.
|
||||
// - 32bytes * 16 registers = 512 bytes for YMM registers.
|
||||
// There are known game failures on real x86 hardware where a 32-bit game is running up against the wall on stack space on non-AVX
|
||||
// hardware and then explodes when run on AVX hardware. This is to guard against that.
|
||||
SupportsAVX = false;
|
||||
}
|
||||
|
||||
OverrideFeatures(this, ForceSVEWidth());
|
||||
}
|
||||
} // namespace FEXCore
|
||||
@@ -185,22 +185,22 @@ bool InterpreterOps::GetFallbackHandler(bool SupportsPreserveAllABI, const IR::I
|
||||
break;
|
||||
}
|
||||
|
||||
#define COMMON_UNARY_X87_OP(OP) \
|
||||
case IR::OP_F80##OP: { \
|
||||
#define COMMON_UNARY_X87_OP(OP) \
|
||||
case IR::OP_F80##OP: { \
|
||||
*Info = {FABI_F80_I16_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80##OP>::handle, Core::OPINDEX_F80##OP, SupportsPreserveAllABI}; \
|
||||
return true; \
|
||||
return true; \
|
||||
}
|
||||
|
||||
#define COMMON_BINARY_X87_OP(OP) \
|
||||
case IR::OP_F80##OP: { \
|
||||
#define COMMON_BINARY_X87_OP(OP) \
|
||||
case IR::OP_F80##OP: { \
|
||||
*Info = {FABI_F80_I16_F80_F80, (void*)&FEXCore::CPU::OpHandlers<IR::OP_F80##OP>::handle, Core::OPINDEX_F80##OP, SupportsPreserveAllABI}; \
|
||||
return true; \
|
||||
return true; \
|
||||
}
|
||||
|
||||
#define COMMON_F64_OP(OP) \
|
||||
case IR::OP_F64##OP: { \
|
||||
#define COMMON_F64_OP(OP) \
|
||||
case IR::OP_F64##OP: { \
|
||||
*Info = GetFallbackInfo(&FEXCore::CPU::OpHandlers<IR::OP_F64##OP>::handle, Core::OPINDEX_F64##OP); \
|
||||
return true; \
|
||||
return true; \
|
||||
}
|
||||
|
||||
// Unary
|
||||
|
||||
@@ -50,22 +50,22 @@ struct OpHandlers<IR::OP_VPCMPESTRX> {
|
||||
|
||||
// Bits are arranged as:
|
||||
// Bit #: 3 2 1 0
|
||||
// [OF | CF | SF | ZF]
|
||||
// [SF | ZF | CF | OF]
|
||||
uint32_t flags = 0;
|
||||
flags |= (valid_rhs < upper_limit) ? 0b01 : 0b00;
|
||||
flags |= (valid_lhs < upper_limit) ? 0b10 : 0b00;
|
||||
flags |= (valid_rhs < upper_limit) ? 0b0100 : 0b0000;
|
||||
flags |= (valid_lhs < upper_limit) ? 0b1000 : 0b0000;
|
||||
|
||||
const uint32_t result = HandlePolarity(aggregation, control, upper_limit, valid_rhs);
|
||||
if (result != 0) {
|
||||
flags |= 0b0100;
|
||||
flags |= 0b0010;
|
||||
}
|
||||
if ((result & 1) != 0) {
|
||||
flags |= 0b1000;
|
||||
flags |= 0b0001;
|
||||
}
|
||||
|
||||
// We tack the flags on top of the result to avoid needing to handle
|
||||
// multiple return values in the JITs.
|
||||
return result | (flags << 16);
|
||||
// We track the flags in the usual NZCV bit position so we can msr them
|
||||
// later. Avoids handling flags natively in JIT.
|
||||
return result | (flags << 28);
|
||||
}
|
||||
|
||||
FEXCORE_PRESERVE_ALL_ATTR static int32_t GetExplicitLength(uint64_t reg, uint16_t control) {
|
||||
|
||||
@@ -7,8 +7,6 @@ $end_info$
|
||||
|
||||
#include "FEXCore/IR/IR.h"
|
||||
#include "Interface/Context/Context.h"
|
||||
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
|
||||
#include "Interface/Core/ArchHelpers/CodeEmitter/Registers.h"
|
||||
#include "Interface/Core/JIT/Arm64/JITClass.h"
|
||||
#include "Interface/IR/Passes/RegisterAllocationPass.h"
|
||||
|
||||
@@ -18,6 +16,31 @@ namespace FEXCore::CPU {
|
||||
#define GRS(Node) (IROp->Size <= 4 ? GetReg<RA_32>(Node) : GetReg<RA_64>(Node))
|
||||
|
||||
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const* IROp, IR::NodeID Node)
|
||||
|
||||
#define DEF_BINOP_WITH_CONSTANT(FEXOp, VarOp, ConstOp) \
|
||||
DEF_OP(FEXOp) { \
|
||||
auto Op = IROp->C<IR::IROp_##FEXOp>(); \
|
||||
\
|
||||
uint64_t Const; \
|
||||
if (IsInlineConstant(Op->Src2, &Const)) { \
|
||||
ConstOp(ConvertSize(IROp), GetReg(Node), GetReg(Op->Src1.ID()), Const); \
|
||||
} else { \
|
||||
VarOp(ConvertSize(IROp), GetReg(Node), GetZeroableReg(Op->Src1), GetReg(Op->Src2.ID())); \
|
||||
} \
|
||||
}
|
||||
|
||||
DEF_BINOP_WITH_CONSTANT(Add, add, add)
|
||||
DEF_BINOP_WITH_CONSTANT(Sub, sub, sub)
|
||||
DEF_BINOP_WITH_CONSTANT(AddWithFlags, adds, adds)
|
||||
DEF_BINOP_WITH_CONSTANT(SubWithFlags, subs, subs)
|
||||
DEF_BINOP_WITH_CONSTANT(Or, orr, orr)
|
||||
DEF_BINOP_WITH_CONSTANT(And, and_, and_)
|
||||
DEF_BINOP_WITH_CONSTANT(Andn, bic, bic)
|
||||
DEF_BINOP_WITH_CONSTANT(Xor, eor, eor)
|
||||
DEF_BINOP_WITH_CONSTANT(Lshl, lslv, lsl)
|
||||
DEF_BINOP_WITH_CONSTANT(Lshr, lsrv, lsr)
|
||||
DEF_BINOP_WITH_CONSTANT(Ror, rorv, ror)
|
||||
|
||||
DEF_OP(TruncElementPair) {
|
||||
auto Op = IROp->C<IR::IROp_TruncElementPair>();
|
||||
|
||||
@@ -69,133 +92,71 @@ DEF_OP(CycleCounter) {
|
||||
#endif
|
||||
}
|
||||
|
||||
DEF_OP(Add) {
|
||||
auto Op = IROp->C<IR::IROp_Add>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
uint64_t Const;
|
||||
if (IsInlineConstant(Op->Src2, &Const)) {
|
||||
add(EmitSize, GetReg(Node), GetReg(Op->Src1.ID()), Const);
|
||||
} else {
|
||||
add(EmitSize, GetReg(Node), GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()));
|
||||
}
|
||||
}
|
||||
|
||||
DEF_OP(AddWithFlags) {
|
||||
auto Op = IROp->C<IR::IROp_AddWithFlags>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
|
||||
const auto EmitSize = OpSize == IR::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
uint64_t Const;
|
||||
if (IsInlineConstant(Op->Src2, &Const)) {
|
||||
adds(EmitSize, GetReg(Node), GetReg(Op->Src1.ID()), Const);
|
||||
} else {
|
||||
adds(EmitSize, GetReg(Node), GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()));
|
||||
}
|
||||
}
|
||||
|
||||
DEF_OP(AddShift) {
|
||||
auto Op = IROp->C<IR::IROp_AddShift>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
add(EmitSize, GetReg(Node), GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()), ConvertIRShiftType(Op->Shift), Op->ShiftAmount);
|
||||
add(ConvertSize48(IROp), GetReg(Node), GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()), ConvertIRShiftType(Op->Shift), Op->ShiftAmount);
|
||||
}
|
||||
|
||||
DEF_OP(AddNZCV) {
|
||||
auto Op = IROp->C<IR::IROp_AddNZCV>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
|
||||
const auto EmitSize = OpSize == IR::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
auto Src1 = GetReg(Op->Src1.ID());
|
||||
|
||||
uint64_t Const;
|
||||
if (IsInlineConstant(Op->Src2, &Const)) {
|
||||
LOGMAN_THROW_AA_FMT(OpSize >= 4, "Constant not allowed here");
|
||||
LOGMAN_THROW_AA_FMT(IROp->Size >= 4, "Constant not allowed here");
|
||||
cmn(EmitSize, Src1, Const);
|
||||
} else {
|
||||
unsigned Shift = OpSize < 4 ? (32 - (8 * OpSize)) : 0;
|
||||
} else if (IROp->Size < 4) {
|
||||
unsigned Shift = 32 - (8 * IROp->Size);
|
||||
|
||||
if (OpSize < 4) {
|
||||
lsl(ARMEmitter::Size::i32Bit, TMP1, Src1, Shift);
|
||||
cmn(EmitSize, TMP1, GetReg(Op->Src2.ID()), ARMEmitter::ShiftType::LSL, Shift);
|
||||
} else {
|
||||
cmn(EmitSize, Src1, GetReg(Op->Src2.ID()));
|
||||
}
|
||||
lsl(ARMEmitter::Size::i32Bit, TMP1, Src1, Shift);
|
||||
cmn(EmitSize, TMP1, GetReg(Op->Src2.ID()), ARMEmitter::ShiftType::LSL, Shift);
|
||||
} else {
|
||||
cmn(EmitSize, Src1, GetReg(Op->Src2.ID()));
|
||||
}
|
||||
}
|
||||
|
||||
DEF_OP(AdcNZCV) {
|
||||
auto Op = IROp->C<IR::IROp_AdcNZCV>();
|
||||
const auto OpSize = IROp->Size;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(OpSize == IR::i32Bit || OpSize == IR::i64Bit, "Unsupported {} size: {}", __func__, OpSize);
|
||||
const auto EmitSize = OpSize == IR::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
adcs(EmitSize, ARMEmitter::Reg::zr, GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()));
|
||||
adcs(ConvertSize48(IROp), ARMEmitter::Reg::zr, GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()));
|
||||
}
|
||||
|
||||
DEF_OP(AdcWithFlags) {
|
||||
auto Op = IROp->C<IR::IROp_AdcWithFlags>();
|
||||
const auto OpSize = IROp->Size;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(OpSize == IR::i32Bit || OpSize == IR::i64Bit, "Unsupported {} size: {}", __func__, OpSize);
|
||||
const auto EmitSize = OpSize == IR::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
adcs(EmitSize, GetReg(Node), GetZeroableReg(Op->Src1), GetReg(Op->Src2.ID()));
|
||||
adcs(ConvertSize48(IROp), GetReg(Node), GetZeroableReg(Op->Src1), GetReg(Op->Src2.ID()));
|
||||
}
|
||||
|
||||
DEF_OP(Adc) {
|
||||
auto Op = IROp->C<IR::IROp_Adc>();
|
||||
const auto OpSize = IROp->Size;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(OpSize == IR::i32Bit || OpSize == IR::i64Bit, "Unsupported {} size: {}", __func__, OpSize);
|
||||
const auto EmitSize = OpSize == IR::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
adc(EmitSize, GetReg(Node), GetZeroableReg(Op->Src1), GetReg(Op->Src2.ID()));
|
||||
adc(ConvertSize48(IROp), GetReg(Node), GetZeroableReg(Op->Src1), GetReg(Op->Src2.ID()));
|
||||
}
|
||||
|
||||
DEF_OP(SbbWithFlags) {
|
||||
auto Op = IROp->C<IR::IROp_SbbWithFlags>();
|
||||
const auto OpSize = IROp->Size;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(OpSize == IR::i32Bit || OpSize == IR::i64Bit, "Unsupported {} size: {}", __func__, OpSize);
|
||||
const auto EmitSize = OpSize == IR::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
sbcs(EmitSize, GetReg(Node), GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()));
|
||||
sbcs(ConvertSize48(IROp), GetReg(Node), GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()));
|
||||
}
|
||||
|
||||
DEF_OP(SbbNZCV) {
|
||||
auto Op = IROp->C<IR::IROp_SbbNZCV>();
|
||||
const auto OpSize = IROp->Size;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(OpSize == IR::i32Bit || OpSize == IR::i64Bit, "Unsupported {} size: {}", __func__, OpSize);
|
||||
const auto EmitSize = OpSize == IR::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
sbcs(EmitSize, ARMEmitter::Reg::zr, GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()));
|
||||
sbcs(ConvertSize48(IROp), ARMEmitter::Reg::zr, GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()));
|
||||
}
|
||||
|
||||
DEF_OP(Sbb) {
|
||||
auto Op = IROp->C<IR::IROp_Sbb>();
|
||||
const auto OpSize = IROp->Size;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(OpSize == IR::i32Bit || OpSize == IR::i64Bit, "Unsupported {} size: {}", __func__, OpSize);
|
||||
const auto EmitSize = OpSize == IR::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
sbc(EmitSize, GetReg(Node), GetZeroableReg(Op->Src1), GetReg(Op->Src2.ID()));
|
||||
sbc(ConvertSize48(IROp), GetReg(Node), GetZeroableReg(Op->Src1), GetReg(Op->Src2.ID()));
|
||||
}
|
||||
|
||||
DEF_OP(TestNZ) {
|
||||
auto Op = IROp->C<IR::IROp_TestNZ>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
|
||||
uint64_t Const;
|
||||
auto Src1 = GetReg(Op->Src1.ID());
|
||||
@@ -203,7 +164,7 @@ DEF_OP(TestNZ) {
|
||||
// Shift the sign bit into place, clearing out the garbage in upper bits.
|
||||
// Adding zero does an effective test, setting NZ according to the result and
|
||||
// zeroing CV.
|
||||
if (OpSize < 4) {
|
||||
if (IROp->Size < 4) {
|
||||
// Cheaper to and+cmn than to lsl+lsl+tst, so do the and ourselves if
|
||||
// needed.
|
||||
if (Op->Src1 != Op->Src2) {
|
||||
@@ -217,7 +178,7 @@ DEF_OP(TestNZ) {
|
||||
Src1 = TMP1;
|
||||
}
|
||||
|
||||
unsigned Shift = 32 - (OpSize * 8);
|
||||
unsigned Shift = 32 - (IROp->Size * 8);
|
||||
cmn(EmitSize, ARMEmitter::Reg::zr, Src1, ARMEmitter::ShiftType::LSL, Shift);
|
||||
} else {
|
||||
if (IsInlineConstant(Op->Src2, &Const)) {
|
||||
@@ -229,51 +190,16 @@ DEF_OP(TestNZ) {
|
||||
}
|
||||
}
|
||||
|
||||
DEF_OP(Sub) {
|
||||
auto Op = IROp->C<IR::IROp_Sub>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
uint64_t Const;
|
||||
if (IsInlineConstant(Op->Src2, &Const)) {
|
||||
sub(EmitSize, GetReg(Node), GetReg(Op->Src1.ID()), Const);
|
||||
} else {
|
||||
sub(EmitSize, GetReg(Node), GetZeroableReg(Op->Src1), GetReg(Op->Src2.ID()));
|
||||
}
|
||||
}
|
||||
|
||||
DEF_OP(SubShift) {
|
||||
auto Op = IROp->C<IR::IROp_SubShift>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
sub(EmitSize, GetReg(Node), GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()), ConvertIRShiftType(Op->Shift), Op->ShiftAmount);
|
||||
}
|
||||
|
||||
DEF_OP(SubWithFlags) {
|
||||
auto Op = IROp->C<IR::IROp_SubWithFlags>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
|
||||
const auto EmitSize = OpSize == IR::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
uint64_t Const;
|
||||
if (IsInlineConstant(Op->Src2, &Const)) {
|
||||
subs(EmitSize, GetReg(Node), GetZeroableReg(Op->Src1), Const);
|
||||
} else {
|
||||
subs(EmitSize, GetReg(Node), GetZeroableReg(Op->Src1), GetReg(Op->Src2.ID()));
|
||||
}
|
||||
sub(ConvertSize48(IROp), GetReg(Node), GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()), ConvertIRShiftType(Op->Shift), Op->ShiftAmount);
|
||||
}
|
||||
|
||||
DEF_OP(SubNZCV) {
|
||||
auto Op = IROp->C<IR::IROp_SubNZCV>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
|
||||
const auto EmitSize = OpSize == IR::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
|
||||
uint64_t Const;
|
||||
if (IsInlineConstant(Op->Src2, &Const)) {
|
||||
@@ -300,9 +226,7 @@ DEF_OP(SubNZCV) {
|
||||
|
||||
DEF_OP(CmpPairZ) {
|
||||
auto Op = IROp->C<IR::IROp_CmpPairZ>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
|
||||
const auto EmitSize = OpSize == IR::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
|
||||
// Save NZCV
|
||||
mrs(TMP1, ARMEmitter::SystemRegister::NZCV);
|
||||
@@ -354,100 +278,54 @@ DEF_OP(AXFlag) {
|
||||
axflag();
|
||||
}
|
||||
|
||||
ARMEmitter::Condition MapSelectCC(IR::CondClassType Cond) {
|
||||
switch (Cond.Val) {
|
||||
case FEXCore::IR::COND_EQ: return ARMEmitter::Condition::CC_EQ;
|
||||
case FEXCore::IR::COND_NEQ: return ARMEmitter::Condition::CC_NE;
|
||||
case FEXCore::IR::COND_SGE: return ARMEmitter::Condition::CC_GE;
|
||||
case FEXCore::IR::COND_SLT: return ARMEmitter::Condition::CC_LT;
|
||||
case FEXCore::IR::COND_SGT: return ARMEmitter::Condition::CC_GT;
|
||||
case FEXCore::IR::COND_SLE: return ARMEmitter::Condition::CC_LE;
|
||||
case FEXCore::IR::COND_UGE: return ARMEmitter::Condition::CC_CS;
|
||||
case FEXCore::IR::COND_ULT: return ARMEmitter::Condition::CC_CC;
|
||||
case FEXCore::IR::COND_UGT: return ARMEmitter::Condition::CC_HI;
|
||||
case FEXCore::IR::COND_ULE: return ARMEmitter::Condition::CC_LS;
|
||||
case FEXCore::IR::COND_FLU: return ARMEmitter::Condition::CC_LT;
|
||||
case FEXCore::IR::COND_FGE: return ARMEmitter::Condition::CC_GE;
|
||||
case FEXCore::IR::COND_FLEU: return ARMEmitter::Condition::CC_LE;
|
||||
case FEXCore::IR::COND_FGT: return ARMEmitter::Condition::CC_GT;
|
||||
case FEXCore::IR::COND_FU: return ARMEmitter::Condition::CC_VS;
|
||||
case FEXCore::IR::COND_FNU: return ARMEmitter::Condition::CC_VC;
|
||||
case FEXCore::IR::COND_VS:
|
||||
case FEXCore::IR::COND_VC:
|
||||
case FEXCore::IR::COND_MI: return ARMEmitter::Condition::CC_MI;
|
||||
case FEXCore::IR::COND_PL: return ARMEmitter::Condition::CC_PL;
|
||||
default: LOGMAN_MSG_A_FMT("Unsupported compare type"); return ARMEmitter::Condition::CC_NV;
|
||||
}
|
||||
}
|
||||
|
||||
DEF_OP(CondAddNZCV) {
|
||||
auto Op = IROp->C<IR::IROp_CondAddNZCV>();
|
||||
const auto OpSize = IROp->Size;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(OpSize == IR::i32Bit || OpSize == IR::i64Bit, "Unsupported {} size: {}", __func__, OpSize);
|
||||
const auto EmitSize = OpSize == IR::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
ARMEmitter::StatusFlags Flags = (ARMEmitter::StatusFlags)Op->FalseNZCV;
|
||||
uint64_t Const = 0;
|
||||
auto Src1 = GetZeroableReg(Op->Src1);
|
||||
|
||||
if (IsInlineConstant(Op->Src2, &Const)) {
|
||||
ccmn(EmitSize, Src1, Const, Flags, MapSelectCC(Op->Cond));
|
||||
ccmn(ConvertSize48(IROp), Src1, Const, Flags, MapCC(Op->Cond));
|
||||
} else {
|
||||
ccmn(EmitSize, Src1, GetReg(Op->Src2.ID()), Flags, MapSelectCC(Op->Cond));
|
||||
ccmn(ConvertSize48(IROp), Src1, GetReg(Op->Src2.ID()), Flags, MapCC(Op->Cond));
|
||||
}
|
||||
}
|
||||
|
||||
DEF_OP(CondSubNZCV) {
|
||||
auto Op = IROp->C<IR::IROp_CondSubNZCV>();
|
||||
const auto OpSize = IROp->Size;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(OpSize == IR::i32Bit || OpSize == IR::i64Bit, "Unsupported {} size: {}", __func__, OpSize);
|
||||
const auto EmitSize = OpSize == IR::i64Bit ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
ARMEmitter::StatusFlags Flags = (ARMEmitter::StatusFlags)Op->FalseNZCV;
|
||||
uint64_t Const = 0;
|
||||
auto Src1 = GetZeroableReg(Op->Src1);
|
||||
|
||||
if (IsInlineConstant(Op->Src2, &Const)) {
|
||||
ccmp(EmitSize, Src1, Const, Flags, MapSelectCC(Op->Cond));
|
||||
ccmp(ConvertSize48(IROp), Src1, Const, Flags, MapCC(Op->Cond));
|
||||
} else {
|
||||
ccmp(EmitSize, Src1, GetReg(Op->Src2.ID()), Flags, MapSelectCC(Op->Cond));
|
||||
ccmp(ConvertSize48(IROp), Src1, GetReg(Op->Src2.ID()), Flags, MapCC(Op->Cond));
|
||||
}
|
||||
}
|
||||
|
||||
DEF_OP(Neg) {
|
||||
auto Op = IROp->C<IR::IROp_Neg>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
if (Op->Cond == FEXCore::IR::COND_AL) {
|
||||
neg(EmitSize, GetReg(Node), GetReg(Op->Src.ID()));
|
||||
neg(ConvertSize48(IROp), GetReg(Node), GetReg(Op->Src.ID()));
|
||||
} else {
|
||||
cneg(EmitSize, GetReg(Node), GetReg(Op->Src.ID()), MapSelectCC(Op->Cond));
|
||||
cneg(ConvertSize48(IROp), GetReg(Node), GetReg(Op->Src.ID()), MapCC(Op->Cond));
|
||||
}
|
||||
}
|
||||
|
||||
DEF_OP(Mul) {
|
||||
auto Op = IROp->C<IR::IROp_Mul>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
mul(EmitSize, GetReg(Node), GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()));
|
||||
mul(ConvertSize48(IROp), GetReg(Node), GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()));
|
||||
}
|
||||
|
||||
DEF_OP(UMul) {
|
||||
auto Op = IROp->C<IR::IROp_UMul>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
mul(EmitSize, GetReg(Node), GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()));
|
||||
mul(ConvertSize48(IROp), GetReg(Node), GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()));
|
||||
}
|
||||
|
||||
DEF_OP(UMull) {
|
||||
@@ -466,13 +344,12 @@ DEF_OP(Div) {
|
||||
// Each source is OpSize in size
|
||||
// So you can have up to a 128bit divide from x86-64
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
|
||||
const auto Dst = GetReg(Node);
|
||||
auto Src1 = GetReg(Op->Src1.ID());
|
||||
auto Src2 = GetReg(Op->Src2.ID());
|
||||
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
if (OpSize == 1) {
|
||||
sxtb(EmitSize, TMP1, Src1);
|
||||
sxtb(EmitSize, TMP2, Src2);
|
||||
@@ -496,13 +373,12 @@ DEF_OP(UDiv) {
|
||||
// Each source is OpSize in size
|
||||
// So you can have up to a 128bit divide from x86-64
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
|
||||
const auto Dst = GetReg(Node);
|
||||
auto Src1 = GetReg(Op->Src1.ID());
|
||||
auto Src2 = GetReg(Op->Src2.ID());
|
||||
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
if (OpSize == 1) {
|
||||
uxtb(EmitSize, TMP1, Src1);
|
||||
uxtb(EmitSize, TMP2, Src2);
|
||||
@@ -525,13 +401,12 @@ DEF_OP(Rem) {
|
||||
// Each source is OpSize in size
|
||||
// So you can have up to a 128bit divide from x86-64
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
|
||||
const auto Dst = GetReg(Node);
|
||||
auto Src1 = GetReg(Op->Src1.ID());
|
||||
auto Src2 = GetReg(Op->Src2.ID());
|
||||
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
if (OpSize == 1) {
|
||||
sxtb(EmitSize, TMP1, Src1);
|
||||
sxtb(EmitSize, TMP2, Src2);
|
||||
@@ -555,12 +430,12 @@ DEF_OP(URem) {
|
||||
// Each source is OpSize in size
|
||||
// So you can have up to a 128bit divide from x86-64
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
|
||||
const auto Dst = GetReg(Node);
|
||||
auto Src1 = GetReg(Op->Src1.ID());
|
||||
auto Src2 = GetReg(Op->Src2.ID());
|
||||
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
if (OpSize == 1) {
|
||||
uxtb(EmitSize, TMP1, Src1);
|
||||
uxtb(EmitSize, TMP2, Src2);
|
||||
@@ -619,90 +494,49 @@ DEF_OP(UMulH) {
|
||||
}
|
||||
}
|
||||
|
||||
DEF_OP(Or) {
|
||||
auto Op = IROp->C<IR::IROp_Or>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
const auto Dst = GetReg(Node);
|
||||
const auto Src1 = GetReg(Op->Src1.ID());
|
||||
|
||||
uint64_t Const;
|
||||
if (IsInlineConstant(Op->Src2, &Const)) {
|
||||
orr(EmitSize, Dst, Src1, Const);
|
||||
} else {
|
||||
const auto Src2 = GetReg(Op->Src2.ID());
|
||||
orr(EmitSize, Dst, Src1, Src2);
|
||||
}
|
||||
}
|
||||
|
||||
DEF_OP(Orlshl) {
|
||||
auto Op = IROp->C<IR::IROp_Orlshl>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
const auto Dst = GetReg(Node);
|
||||
const auto Src1 = GetReg(Op->Src1.ID());
|
||||
|
||||
uint64_t Const;
|
||||
if (IsInlineConstant(Op->Src2, &Const)) {
|
||||
orr(EmitSize, Dst, Src1, Const << Op->BitShift);
|
||||
orr(ConvertSize(IROp), Dst, Src1, Const << Op->BitShift);
|
||||
} else {
|
||||
const auto Src2 = GetReg(Op->Src2.ID());
|
||||
orr(EmitSize, Dst, Src1, Src2, ARMEmitter::ShiftType::LSL, Op->BitShift);
|
||||
orr(ConvertSize(IROp), Dst, Src1, Src2, ARMEmitter::ShiftType::LSL, Op->BitShift);
|
||||
}
|
||||
}
|
||||
|
||||
DEF_OP(Orlshr) {
|
||||
auto Op = IROp->C<IR::IROp_Orlshr>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
const auto Dst = GetReg(Node);
|
||||
const auto Src1 = GetReg(Op->Src1.ID());
|
||||
|
||||
uint64_t Const;
|
||||
if (IsInlineConstant(Op->Src2, &Const)) {
|
||||
orr(EmitSize, Dst, Src1, Const >> Op->BitShift);
|
||||
orr(ConvertSize(IROp), Dst, Src1, Const >> Op->BitShift);
|
||||
} else {
|
||||
const auto Src2 = GetReg(Op->Src2.ID());
|
||||
orr(EmitSize, Dst, Src1, Src2, ARMEmitter::ShiftType::LSR, Op->BitShift);
|
||||
orr(ConvertSize(IROp), Dst, Src1, Src2, ARMEmitter::ShiftType::LSR, Op->BitShift);
|
||||
}
|
||||
}
|
||||
|
||||
DEF_OP(Ornror) {
|
||||
auto Op = IROp->C<IR::IROp_Ornror>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
const auto Dst = GetReg(Node);
|
||||
const auto Src1 = GetReg(Op->Src1.ID());
|
||||
|
||||
const auto Src2 = GetReg(Op->Src2.ID());
|
||||
orn(EmitSize, Dst, Src1, Src2, ARMEmitter::ShiftType::ROR, Op->BitShift);
|
||||
}
|
||||
|
||||
DEF_OP(And) {
|
||||
auto Op = IROp->C<IR::IROp_And>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
const auto Dst = GetReg(Node);
|
||||
const auto Src1 = GetReg(Op->Src1.ID());
|
||||
|
||||
uint64_t Const;
|
||||
if (IsInlineConstant(Op->Src2, &Const)) {
|
||||
and_(EmitSize, Dst, Src1, Const);
|
||||
} else {
|
||||
const auto Src2 = GetReg(Op->Src2.ID());
|
||||
and_(EmitSize, Dst, Src1, Src2);
|
||||
}
|
||||
orn(ConvertSize(IROp), Dst, Src1, Src2, ARMEmitter::ShiftType::ROR, Op->BitShift);
|
||||
}
|
||||
|
||||
DEF_OP(AndWithFlags) {
|
||||
auto Op = IROp->C<IR::IROp_AndWithFlags>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
|
||||
uint64_t Const;
|
||||
const auto Dst = GetReg(Node);
|
||||
@@ -734,99 +568,22 @@ DEF_OP(AndWithFlags) {
|
||||
}
|
||||
}
|
||||
|
||||
DEF_OP(Andn) {
|
||||
auto Op = IROp->C<IR::IROp_Andn>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
const auto Dst = GetReg(Node);
|
||||
const auto Src1 = GetReg(Op->Src1.ID());
|
||||
|
||||
uint64_t Const;
|
||||
if (IsInlineConstant(Op->Src2, &Const)) {
|
||||
bic(EmitSize, Dst, Src1, Const);
|
||||
} else {
|
||||
const auto Src2 = GetReg(Op->Src2.ID());
|
||||
bic(EmitSize, Dst, Src1, Src2);
|
||||
}
|
||||
}
|
||||
|
||||
DEF_OP(Xor) {
|
||||
auto Op = IROp->C<IR::IROp_Xor>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
const auto Dst = GetReg(Node);
|
||||
const auto Src1 = GetReg(Op->Src1.ID());
|
||||
|
||||
uint64_t Const;
|
||||
if (IsInlineConstant(Op->Src2, &Const)) {
|
||||
eor(EmitSize, Dst, Src1, Const);
|
||||
} else {
|
||||
const auto Src2 = GetReg(Op->Src2.ID());
|
||||
eor(EmitSize, Dst, Src1, Src2);
|
||||
}
|
||||
}
|
||||
|
||||
DEF_OP(XorShift) {
|
||||
auto Op = IROp->C<IR::IROp_XorShift>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
eor(EmitSize, GetReg(Node), GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()), ConvertIRShiftType(Op->Shift), Op->ShiftAmount);
|
||||
eor(ConvertSize48(IROp), GetReg(Node), GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()), ConvertIRShiftType(Op->Shift), Op->ShiftAmount);
|
||||
}
|
||||
|
||||
DEF_OP(XornShift) {
|
||||
auto Op = IROp->C<IR::IROp_XornShift>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
eon(EmitSize, GetReg(Node), GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()), ConvertIRShiftType(Op->Shift), Op->ShiftAmount);
|
||||
}
|
||||
|
||||
DEF_OP(Lshl) {
|
||||
auto Op = IROp->C<IR::IROp_Lshl>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
const auto Dst = GetReg(Node);
|
||||
const auto Src1 = GetReg(Op->Src1.ID());
|
||||
|
||||
uint64_t Const;
|
||||
if (IsInlineConstant(Op->Src2, &Const)) {
|
||||
lsl(EmitSize, Dst, Src1, Const);
|
||||
} else {
|
||||
const auto Src2 = GetReg(Op->Src2.ID());
|
||||
lslv(EmitSize, Dst, Src1, Src2);
|
||||
}
|
||||
}
|
||||
|
||||
DEF_OP(Lshr) {
|
||||
auto Op = IROp->C<IR::IROp_Lshr>();
|
||||
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
const auto Dst = GetReg(Node);
|
||||
const auto Src1 = GetReg(Op->Src1.ID());
|
||||
|
||||
uint64_t Const;
|
||||
if (IsInlineConstant(Op->Src2, &Const)) {
|
||||
lsr(EmitSize, Dst, Src1, Const);
|
||||
} else {
|
||||
const auto Src2 = GetReg(Op->Src2.ID());
|
||||
lsrv(EmitSize, Dst, Src1, Src2);
|
||||
}
|
||||
eon(ConvertSize48(IROp), GetReg(Node), GetReg(Op->Src1.ID()), GetReg(Op->Src2.ID()), ConvertIRShiftType(Op->Shift), Op->ShiftAmount);
|
||||
}
|
||||
|
||||
DEF_OP(Ashr) {
|
||||
auto Op = IROp->C<IR::IROp_Ashr>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
|
||||
const auto Dst = GetReg(Node);
|
||||
const auto Src1 = GetReg(Op->Src1.ID());
|
||||
@@ -932,45 +689,18 @@ DEF_OP(ShiftFlags) {
|
||||
}
|
||||
}
|
||||
|
||||
DEF_OP(Ror) {
|
||||
auto Op = IROp->C<IR::IROp_Ror>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
const auto Dst = GetReg(Node);
|
||||
const auto Src1 = GetReg(Op->Src1.ID());
|
||||
|
||||
uint64_t Const;
|
||||
if (IsInlineConstant(Op->Src2, &Const)) {
|
||||
ror(EmitSize, Dst, Src1, Const);
|
||||
} else {
|
||||
const auto Src2 = GetReg(Op->Src2.ID());
|
||||
rorv(EmitSize, Dst, Src1, Src2);
|
||||
}
|
||||
}
|
||||
|
||||
DEF_OP(Extr) {
|
||||
auto Op = IROp->C<IR::IROp_Extr>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
const auto Dst = GetReg(Node);
|
||||
const auto Upper = GetReg(Op->Upper.ID());
|
||||
const auto Lower = GetReg(Op->Lower.ID());
|
||||
|
||||
extr(EmitSize, Dst, Upper, Lower, Op->LSB);
|
||||
extr(ConvertSize48(IROp), Dst, Upper, Lower, Op->LSB);
|
||||
}
|
||||
|
||||
DEF_OP(PDep) {
|
||||
auto Op = IROp->C<IR::IROp_PExt>();
|
||||
const auto OpSize = IROp->Size;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto EmitSize = ConvertSize48(IROp);
|
||||
|
||||
const auto Dest = GetReg(Node);
|
||||
|
||||
@@ -1033,9 +763,7 @@ DEF_OP(PExt) {
|
||||
auto Op = IROp->C<IR::IROp_PExt>();
|
||||
const auto OpSize = IROp->Size;
|
||||
const auto OpSizeBitsM1 = (OpSize * 8) - 1;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto EmitSize = ConvertSize48(IROp);
|
||||
|
||||
const auto Input = GetReg(Op->Input.ID());
|
||||
const auto Mask = GetReg(Op->Mask.ID());
|
||||
@@ -1351,15 +1079,11 @@ DEF_OP(LURem) {
|
||||
|
||||
DEF_OP(Not) {
|
||||
auto Op = IROp->C<IR::IROp_Not>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
const auto Dst = GetReg(Node);
|
||||
const auto Src = GetReg(Op->Src.ID());
|
||||
|
||||
mvn(EmitSize, Dst, Src);
|
||||
mvn(ConvertSize48(IROp), Dst, Src);
|
||||
}
|
||||
|
||||
DEF_OP(Popcount) {
|
||||
@@ -1373,23 +1097,23 @@ DEF_OP(Popcount) {
|
||||
case 0x1:
|
||||
fmov(ARMEmitter::Size::i32Bit, VTMP1.S(), Src);
|
||||
// only use lowest byte
|
||||
cnt(FEXCore::ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D());
|
||||
cnt(ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D());
|
||||
break;
|
||||
case 0x2:
|
||||
fmov(ARMEmitter::Size::i32Bit, VTMP1.S(), Src);
|
||||
cnt(FEXCore::ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D());
|
||||
cnt(ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D());
|
||||
// only count two lowest bytes
|
||||
addp(FEXCore::ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D(), VTMP1.D());
|
||||
addp(ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D(), VTMP1.D());
|
||||
break;
|
||||
case 0x4:
|
||||
fmov(ARMEmitter::Size::i32Bit, VTMP1.S(), Src);
|
||||
cnt(FEXCore::ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D());
|
||||
cnt(ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D());
|
||||
// fmov has zero extended, unused bytes are zero
|
||||
addv(ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D());
|
||||
break;
|
||||
case 0x8:
|
||||
fmov(ARMEmitter::Size::i64Bit, VTMP1.D(), Src);
|
||||
cnt(FEXCore::ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D());
|
||||
cnt(ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D());
|
||||
// fmov has zero extended, unused bytes are zero
|
||||
addv(ARMEmitter::SubRegSize::i8Bit, VTMP1.D(), VTMP1.D());
|
||||
break;
|
||||
@@ -1401,15 +1125,13 @@ DEF_OP(Popcount) {
|
||||
|
||||
DEF_OP(FindLSB) {
|
||||
auto Op = IROp->C<IR::IROp_FindLSB>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
|
||||
const auto Dst = GetReg(Node);
|
||||
const auto Src = GetReg(Op->Src.ID());
|
||||
|
||||
if (OpSize != 8) {
|
||||
ubfx(EmitSize, TMP1, Src, 0, OpSize * 8);
|
||||
if (IROp->Size != 8) {
|
||||
ubfx(EmitSize, TMP1, Src, 0, IROp->Size * 8);
|
||||
cmp(EmitSize, TMP1, 0);
|
||||
rbit(EmitSize, TMP1, TMP1);
|
||||
} else {
|
||||
@@ -1426,7 +1148,7 @@ DEF_OP(FindMSB) {
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 2 || OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
|
||||
const auto Dst = GetReg(Node);
|
||||
const auto Src = GetReg(Op->Src.ID());
|
||||
@@ -1449,7 +1171,7 @@ DEF_OP(FindTrailingZeroes) {
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 2 || OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
|
||||
const auto Dst = GetReg(Node);
|
||||
const auto Src = GetReg(Op->Src.ID());
|
||||
@@ -1473,7 +1195,7 @@ DEF_OP(CountLeadingZeroes) {
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 2 || OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
|
||||
const auto Dst = GetReg(Node);
|
||||
const auto Src = GetReg(Op->Src.ID());
|
||||
@@ -1494,7 +1216,7 @@ DEF_OP(Rev) {
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 2 || OpSize == 4 || OpSize == 8, "Unsupported {} size: {}", __func__, OpSize);
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
|
||||
const auto Dst = GetReg(Node);
|
||||
const auto Src = GetReg(Op->Src.ID());
|
||||
@@ -1507,9 +1229,7 @@ DEF_OP(Rev) {
|
||||
|
||||
DEF_OP(Bfi) {
|
||||
auto Op = IROp->C<IR::IROp_Bfi>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
|
||||
const auto Dst = GetReg(Node);
|
||||
const auto SrcDst = GetReg(Op->Dest.ID());
|
||||
@@ -1528,19 +1248,17 @@ DEF_OP(Bfi) {
|
||||
mov(EmitSize, TMP1, SrcDst);
|
||||
bfi(EmitSize, TMP1, Src, Op->lsb, Op->Width);
|
||||
|
||||
if (OpSize >= 4) {
|
||||
if (IROp->Size >= 4) {
|
||||
mov(EmitSize, Dst, TMP1.R());
|
||||
} else {
|
||||
ubfx(EmitSize, Dst, TMP1, 0, OpSize * 8);
|
||||
ubfx(EmitSize, Dst, TMP1, 0, IROp->Size * 8);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
DEF_OP(Bfxil) {
|
||||
auto Op = IROp->C<IR::IROp_Bfxil>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
|
||||
const auto Dst = GetReg(Node);
|
||||
const auto SrcDst = GetReg(Op->Dest.ID());
|
||||
@@ -1566,8 +1284,7 @@ DEF_OP(Bfe) {
|
||||
auto Op = IROp->C<IR::IROp_Bfe>();
|
||||
LOGMAN_THROW_AA_FMT(IROp->Size <= 8, "OpSize is too large for BFE: {}", IROp->Size);
|
||||
LOGMAN_THROW_AA_FMT(Op->Width != 0, "Invalid BFE width of 0");
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
|
||||
const auto Dst = GetReg(Node);
|
||||
const auto Src = GetReg(Op->Src.ID());
|
||||
@@ -1575,7 +1292,7 @@ DEF_OP(Bfe) {
|
||||
if (Op->lsb == 0 && Op->Width == 32) {
|
||||
mov(ARMEmitter::Size::i32Bit, Dst, Src);
|
||||
} else if (Op->lsb == 0 && Op->Width == 64) {
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 8, "Must be 64-bit wide register");
|
||||
LOGMAN_THROW_AA_FMT(IROp->Size == 8, "Must be 64-bit wide register");
|
||||
mov(ARMEmitter::Size::i64Bit, Dst, Src);
|
||||
} else {
|
||||
ubfx(EmitSize, Dst, Src, Op->lsb, Op->Width);
|
||||
@@ -1584,23 +1301,20 @@ DEF_OP(Bfe) {
|
||||
|
||||
DEF_OP(Sbfe) {
|
||||
auto Op = IROp->C<IR::IROp_Sbfe>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
const auto Dst = GetReg(Node);
|
||||
const auto Src = GetReg(Op->Src.ID());
|
||||
|
||||
sbfx(EmitSize, Dst, Src, Op->lsb, Op->Width);
|
||||
sbfx(ConvertSize(IROp), Dst, Src, Op->lsb, Op->Width);
|
||||
}
|
||||
|
||||
DEF_OP(Select) {
|
||||
auto Op = IROp->C<IR::IROp_Select>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
const auto CompareEmitSize = Op->CompareSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
uint64_t Const;
|
||||
auto cc = MapSelectCC(Op->Cond);
|
||||
auto cc = MapCC(Op->Cond);
|
||||
|
||||
if (IsGPR(Op->Cmp1.ID())) {
|
||||
const auto Src1 = GetReg(Op->Cmp1.ID());
|
||||
@@ -1649,16 +1363,15 @@ DEF_OP(Select) {
|
||||
|
||||
DEF_OP(NZCVSelect) {
|
||||
auto Op = IROp->C<IR::IROp_NZCVSelect>();
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
|
||||
auto cc = MapSelectCC(Op->Cond);
|
||||
auto cc = MapCC(Op->Cond);
|
||||
|
||||
uint64_t const_true, const_false;
|
||||
bool is_const_true = IsInlineConstant(Op->TrueVal, &const_true);
|
||||
bool is_const_false = IsInlineConstant(Op->FalseVal, &const_false);
|
||||
|
||||
uint64_t all_ones = OpSize == 8 ? 0xffff'ffff'ffff'ffffull : 0xffff'ffffull;
|
||||
uint64_t all_ones = IROp->Size == 8 ? 0xffff'ffff'ffff'ffffull : 0xffff'ffffull;
|
||||
|
||||
ARMEmitter::Register Dst = GetReg(Node);
|
||||
|
||||
@@ -1740,12 +1453,11 @@ DEF_OP(Float_ToGPR_ZS) {
|
||||
|
||||
ARMEmitter::Register Dst = GetReg(Node);
|
||||
ARMEmitter::VRegister Src = GetVReg(Op->Scalar.ID());
|
||||
const auto DestSize = IROp->Size == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
if (Op->SrcElementSize == 8) {
|
||||
fcvtzs(DestSize, Dst, Src.D());
|
||||
fcvtzs(ConvertSize(IROp), Dst, Src.D());
|
||||
} else {
|
||||
fcvtzs(DestSize, Dst, Src.S());
|
||||
fcvtzs(ConvertSize(IROp), Dst, Src.S());
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1754,14 +1466,13 @@ DEF_OP(Float_ToGPR_S) {
|
||||
|
||||
ARMEmitter::Register Dst = GetReg(Node);
|
||||
ARMEmitter::VRegister Src = GetVReg(Op->Scalar.ID());
|
||||
const auto DestSize = IROp->Size == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
if (Op->SrcElementSize == 8) {
|
||||
frinti(VTMP1.D(), Src.D());
|
||||
fcvtzs(DestSize, Dst, VTMP1.D());
|
||||
fcvtzs(ConvertSize(IROp), Dst, VTMP1.D());
|
||||
} else {
|
||||
frinti(VTMP1.S(), Src.S());
|
||||
fcvtzs(DestSize, Dst, VTMP1.S());
|
||||
fcvtzs(ConvertSize(IROp), Dst, VTMP1.S());
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -6,7 +6,6 @@ $end_info$
|
||||
*/
|
||||
|
||||
#include "Interface/Context/Context.h"
|
||||
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
|
||||
#include "Interface/Core/Dispatcher/Dispatcher.h"
|
||||
#include "Interface/Core/JIT/Arm64/JITClass.h"
|
||||
|
||||
@@ -69,8 +68,8 @@ DEF_OP(CASPair) {
|
||||
|
||||
DEF_OP(CAS) {
|
||||
auto Op = IROp->C<IR::IROp_CAS>();
|
||||
uint8_t OpSize = IROp->Size;
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
|
||||
// DataSrc = *Src1
|
||||
// if (DataSrc == Src3) { *Src1 == Src2; } Src2 = DataSrc
|
||||
// This will write to memory! Careful!
|
||||
@@ -79,13 +78,6 @@ DEF_OP(CAS) {
|
||||
auto Desired = GetReg(Op->Desired.ID());
|
||||
auto MemSrc = GetReg(Op->Addr.ID());
|
||||
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
|
||||
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
|
||||
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
|
||||
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
|
||||
ARMEmitter::SubRegSize::i8Bit;
|
||||
|
||||
if (CTX->HostFeatures.SupportsAtomics) {
|
||||
mov(EmitSize, TMP2, Expected);
|
||||
casal(SubEmitSize, TMP2, Desired, MemSrc);
|
||||
@@ -96,9 +88,9 @@ DEF_OP(CAS) {
|
||||
ARMEmitter::SingleUseForwardLabel LoopExpected;
|
||||
Bind(&LoopTop);
|
||||
ldaxr(SubEmitSize, TMP2, MemSrc);
|
||||
if (OpSize == 1) {
|
||||
if (IROp->Size == 1) {
|
||||
cmp(EmitSize, TMP2, Expected, ARMEmitter::ExtendedType::UXTB, 0);
|
||||
} else if (OpSize == 2) {
|
||||
} else if (IROp->Size == 2) {
|
||||
cmp(EmitSize, TMP2, Expected, ARMEmitter::ExtendedType::UXTH, 0);
|
||||
} else {
|
||||
cmp(EmitSize, TMP2, Expected);
|
||||
@@ -120,19 +112,12 @@ DEF_OP(CAS) {
|
||||
|
||||
DEF_OP(AtomicAdd) {
|
||||
auto Op = IROp->C<IR::IROp_AtomicAdd>();
|
||||
uint8_t OpSize = IROp->Size;
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
|
||||
|
||||
auto MemSrc = GetReg(Op->Addr.ID());
|
||||
auto Src = GetReg(Op->Value.ID());
|
||||
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
|
||||
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
|
||||
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
|
||||
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
|
||||
ARMEmitter::SubRegSize::i8Bit;
|
||||
|
||||
if (CTX->HostFeatures.SupportsAtomics) {
|
||||
staddl(SubEmitSize, Src, MemSrc);
|
||||
} else {
|
||||
@@ -147,19 +132,12 @@ DEF_OP(AtomicAdd) {
|
||||
|
||||
DEF_OP(AtomicSub) {
|
||||
auto Op = IROp->C<IR::IROp_AtomicSub>();
|
||||
uint8_t OpSize = IROp->Size;
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
|
||||
|
||||
auto MemSrc = GetReg(Op->Addr.ID());
|
||||
auto Src = GetReg(Op->Value.ID());
|
||||
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
|
||||
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
|
||||
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
|
||||
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
|
||||
ARMEmitter::SubRegSize::i8Bit;
|
||||
|
||||
if (CTX->HostFeatures.SupportsAtomics) {
|
||||
neg(EmitSize, TMP2, Src);
|
||||
staddl(SubEmitSize, TMP2, MemSrc);
|
||||
@@ -175,19 +153,12 @@ DEF_OP(AtomicSub) {
|
||||
|
||||
DEF_OP(AtomicAnd) {
|
||||
auto Op = IROp->C<IR::IROp_AtomicAnd>();
|
||||
uint8_t OpSize = IROp->Size;
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
|
||||
|
||||
auto MemSrc = GetReg(Op->Addr.ID());
|
||||
auto Src = GetReg(Op->Value.ID());
|
||||
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
|
||||
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
|
||||
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
|
||||
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
|
||||
ARMEmitter::SubRegSize::i8Bit;
|
||||
|
||||
if (CTX->HostFeatures.SupportsAtomics) {
|
||||
mvn(EmitSize, TMP2, Src);
|
||||
stclrl(SubEmitSize, TMP2, MemSrc);
|
||||
@@ -203,19 +174,12 @@ DEF_OP(AtomicAnd) {
|
||||
|
||||
DEF_OP(AtomicCLR) {
|
||||
auto Op = IROp->C<IR::IROp_AtomicCLR>();
|
||||
uint8_t OpSize = IROp->Size;
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
|
||||
|
||||
auto MemSrc = GetReg(Op->Addr.ID());
|
||||
auto Src = GetReg(Op->Value.ID());
|
||||
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
|
||||
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
|
||||
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
|
||||
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
|
||||
ARMEmitter::SubRegSize::i8Bit;
|
||||
|
||||
if (CTX->HostFeatures.SupportsAtomics) {
|
||||
stclrl(SubEmitSize, Src, MemSrc);
|
||||
} else {
|
||||
@@ -230,19 +194,12 @@ DEF_OP(AtomicCLR) {
|
||||
|
||||
DEF_OP(AtomicOr) {
|
||||
auto Op = IROp->C<IR::IROp_AtomicOr>();
|
||||
uint8_t OpSize = IROp->Size;
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
|
||||
|
||||
auto MemSrc = GetReg(Op->Addr.ID());
|
||||
auto Src = GetReg(Op->Value.ID());
|
||||
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
|
||||
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
|
||||
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
|
||||
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
|
||||
ARMEmitter::SubRegSize::i8Bit;
|
||||
|
||||
if (CTX->HostFeatures.SupportsAtomics) {
|
||||
stsetl(SubEmitSize, Src, MemSrc);
|
||||
} else {
|
||||
@@ -257,19 +214,12 @@ DEF_OP(AtomicOr) {
|
||||
|
||||
DEF_OP(AtomicXor) {
|
||||
auto Op = IROp->C<IR::IROp_AtomicXor>();
|
||||
uint8_t OpSize = IROp->Size;
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
|
||||
|
||||
auto MemSrc = GetReg(Op->Addr.ID());
|
||||
auto Src = GetReg(Op->Value.ID());
|
||||
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
|
||||
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
|
||||
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
|
||||
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
|
||||
ARMEmitter::SubRegSize::i8Bit;
|
||||
|
||||
if (CTX->HostFeatures.SupportsAtomics) {
|
||||
steorl(SubEmitSize, Src, MemSrc);
|
||||
} else {
|
||||
@@ -284,18 +234,11 @@ DEF_OP(AtomicXor) {
|
||||
|
||||
DEF_OP(AtomicNeg) {
|
||||
auto Op = IROp->C<IR::IROp_AtomicNeg>();
|
||||
uint8_t OpSize = IROp->Size;
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
|
||||
|
||||
auto MemSrc = GetReg(Op->Addr.ID());
|
||||
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
|
||||
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
|
||||
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
|
||||
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
|
||||
ARMEmitter::SubRegSize::i8Bit;
|
||||
|
||||
ARMEmitter::BackwardLabel LoopTop;
|
||||
Bind(&LoopTop);
|
||||
ldaxr(SubEmitSize, TMP2, MemSrc);
|
||||
@@ -312,7 +255,7 @@ DEF_OP(AtomicSwap) {
|
||||
auto MemSrc = GetReg(Op->Addr.ID());
|
||||
auto Src = GetReg(Op->Value.ID());
|
||||
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
|
||||
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
|
||||
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
|
||||
@@ -333,19 +276,12 @@ DEF_OP(AtomicSwap) {
|
||||
|
||||
DEF_OP(AtomicFetchAdd) {
|
||||
auto Op = IROp->C<IR::IROp_AtomicFetchAdd>();
|
||||
uint8_t OpSize = IROp->Size;
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
|
||||
|
||||
auto MemSrc = GetReg(Op->Addr.ID());
|
||||
auto Src = GetReg(Op->Value.ID());
|
||||
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
|
||||
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
|
||||
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
|
||||
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
|
||||
ARMEmitter::SubRegSize::i8Bit;
|
||||
|
||||
if (CTX->HostFeatures.SupportsAtomics) {
|
||||
ldaddal(SubEmitSize, Src, GetReg(Node), MemSrc);
|
||||
} else {
|
||||
@@ -361,19 +297,12 @@ DEF_OP(AtomicFetchAdd) {
|
||||
|
||||
DEF_OP(AtomicFetchSub) {
|
||||
auto Op = IROp->C<IR::IROp_AtomicFetchSub>();
|
||||
uint8_t OpSize = IROp->Size;
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
|
||||
|
||||
auto MemSrc = GetReg(Op->Addr.ID());
|
||||
auto Src = GetReg(Op->Value.ID());
|
||||
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
|
||||
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
|
||||
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
|
||||
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
|
||||
ARMEmitter::SubRegSize::i8Bit;
|
||||
|
||||
if (CTX->HostFeatures.SupportsAtomics) {
|
||||
neg(EmitSize, TMP2, Src);
|
||||
ldaddal(SubEmitSize, TMP2, GetReg(Node), MemSrc);
|
||||
@@ -390,19 +319,12 @@ DEF_OP(AtomicFetchSub) {
|
||||
|
||||
DEF_OP(AtomicFetchAnd) {
|
||||
auto Op = IROp->C<IR::IROp_AtomicFetchAnd>();
|
||||
uint8_t OpSize = IROp->Size;
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
|
||||
|
||||
auto MemSrc = GetReg(Op->Addr.ID());
|
||||
auto Src = GetReg(Op->Value.ID());
|
||||
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
|
||||
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
|
||||
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
|
||||
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
|
||||
ARMEmitter::SubRegSize::i8Bit;
|
||||
|
||||
if (CTX->HostFeatures.SupportsAtomics) {
|
||||
mvn(EmitSize, TMP2, Src);
|
||||
ldclral(SubEmitSize, TMP2, GetReg(Node), MemSrc);
|
||||
@@ -419,19 +341,12 @@ DEF_OP(AtomicFetchAnd) {
|
||||
|
||||
DEF_OP(AtomicFetchCLR) {
|
||||
auto Op = IROp->C<IR::IROp_AtomicFetchCLR>();
|
||||
uint8_t OpSize = IROp->Size;
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
|
||||
|
||||
auto MemSrc = GetReg(Op->Addr.ID());
|
||||
auto Src = GetReg(Op->Value.ID());
|
||||
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
|
||||
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
|
||||
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
|
||||
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
|
||||
ARMEmitter::SubRegSize::i8Bit;
|
||||
|
||||
if (CTX->HostFeatures.SupportsAtomics) {
|
||||
ldclral(SubEmitSize, Src, GetReg(Node), MemSrc);
|
||||
} else {
|
||||
@@ -447,19 +362,12 @@ DEF_OP(AtomicFetchCLR) {
|
||||
|
||||
DEF_OP(AtomicFetchOr) {
|
||||
auto Op = IROp->C<IR::IROp_AtomicFetchOr>();
|
||||
uint8_t OpSize = IROp->Size;
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
|
||||
|
||||
auto MemSrc = GetReg(Op->Addr.ID());
|
||||
auto Src = GetReg(Op->Value.ID());
|
||||
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
|
||||
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
|
||||
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
|
||||
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
|
||||
ARMEmitter::SubRegSize::i8Bit;
|
||||
|
||||
if (CTX->HostFeatures.SupportsAtomics) {
|
||||
ldsetal(SubEmitSize, Src, GetReg(Node), MemSrc);
|
||||
} else {
|
||||
@@ -475,19 +383,12 @@ DEF_OP(AtomicFetchOr) {
|
||||
|
||||
DEF_OP(AtomicFetchXor) {
|
||||
auto Op = IROp->C<IR::IROp_AtomicFetchXor>();
|
||||
uint8_t OpSize = IROp->Size;
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
|
||||
|
||||
auto MemSrc = GetReg(Op->Addr.ID());
|
||||
auto Src = GetReg(Op->Value.ID());
|
||||
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
|
||||
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
|
||||
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
|
||||
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
|
||||
ARMEmitter::SubRegSize::i8Bit;
|
||||
|
||||
if (CTX->HostFeatures.SupportsAtomics) {
|
||||
ldeoral(SubEmitSize, Src, GetReg(Node), MemSrc);
|
||||
} else {
|
||||
@@ -503,18 +404,11 @@ DEF_OP(AtomicFetchXor) {
|
||||
|
||||
DEF_OP(AtomicFetchNeg) {
|
||||
auto Op = IROp->C<IR::IROp_AtomicFetchNeg>();
|
||||
uint8_t OpSize = IROp->Size;
|
||||
LOGMAN_THROW_AA_FMT(OpSize == 8 || OpSize == 4 || OpSize == 2 || OpSize == 1, "Unexpected CAS size");
|
||||
const auto EmitSize = ConvertSize(IROp);
|
||||
const auto SubEmitSize = ConvertSubRegSize8(IROp->Size);
|
||||
|
||||
auto MemSrc = GetReg(Op->Addr.ID());
|
||||
|
||||
const auto EmitSize = OpSize == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
const auto SubEmitSize = OpSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
|
||||
OpSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
|
||||
OpSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
|
||||
OpSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
|
||||
ARMEmitter::SubRegSize::i8Bit;
|
||||
|
||||
ARMEmitter::BackwardLabel LoopTop;
|
||||
Bind(&LoopTop);
|
||||
ldaxr(SubEmitSize, TMP2, MemSrc);
|
||||
|
||||
@@ -7,7 +7,6 @@ $end_info$
|
||||
|
||||
#include "Interface/Context/Context.h"
|
||||
#include "FEXCore/IR/IR.h"
|
||||
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
|
||||
#include "Interface/Core/LookupCache.h"
|
||||
|
||||
#include "Interface/Core/JIT/Arm64/JITClass.h"
|
||||
@@ -55,7 +54,8 @@ DEF_OP(ExitFunction) {
|
||||
if (IsInlineConstant(Op->NewRIP, &NewRIP) || IsInlineEntrypointOffset(Op->NewRIP, &NewRIP)) {
|
||||
#ifdef _M_ARM_64EC
|
||||
if (RtlIsEcCode(NewRIP)) {
|
||||
LoadConstant(ARMEmitter::Size::i64Bit, TMP3, NewRIP);
|
||||
add(ARMEmitter::Size::i64Bit, ARMEmitter::Reg::rsp, StaticRegisters[X86State::REG_RSP], 0);
|
||||
LoadConstant(ARMEmitter::Size::i64Bit, EC_CALL_CHECKER_PC_REG, NewRIP);
|
||||
ldr(TMP2, STATE_PTR(CpuStateFrame, Pointers.Common.ExitFunctionEC));
|
||||
br(TMP2);
|
||||
} else {
|
||||
@@ -101,39 +101,13 @@ DEF_OP(Jump) {
|
||||
PendingTargetLabel = &JumpTargets.try_emplace(Target).first->second;
|
||||
}
|
||||
|
||||
static ARMEmitter::Condition MapBranchCC(IR::CondClassType Cond) {
|
||||
switch (Cond.Val) {
|
||||
case FEXCore::IR::COND_EQ: return ARMEmitter::Condition::CC_EQ;
|
||||
case FEXCore::IR::COND_NEQ: return ARMEmitter::Condition::CC_NE;
|
||||
case FEXCore::IR::COND_SGE: return ARMEmitter::Condition::CC_GE;
|
||||
case FEXCore::IR::COND_SLT: return ARMEmitter::Condition::CC_LT;
|
||||
case FEXCore::IR::COND_SGT: return ARMEmitter::Condition::CC_GT;
|
||||
case FEXCore::IR::COND_SLE: return ARMEmitter::Condition::CC_LE;
|
||||
case FEXCore::IR::COND_UGE: return ARMEmitter::Condition::CC_CS;
|
||||
case FEXCore::IR::COND_ULT: return ARMEmitter::Condition::CC_CC;
|
||||
case FEXCore::IR::COND_UGT: return ARMEmitter::Condition::CC_HI;
|
||||
case FEXCore::IR::COND_ULE: return ARMEmitter::Condition::CC_LS;
|
||||
case FEXCore::IR::COND_FLU: return ARMEmitter::Condition::CC_LT;
|
||||
case FEXCore::IR::COND_FGE: return ARMEmitter::Condition::CC_GE;
|
||||
case FEXCore::IR::COND_FLEU: return ARMEmitter::Condition::CC_LE;
|
||||
case FEXCore::IR::COND_FGT: return ARMEmitter::Condition::CC_GT;
|
||||
case FEXCore::IR::COND_FU: return ARMEmitter::Condition::CC_VS;
|
||||
case FEXCore::IR::COND_FNU: return ARMEmitter::Condition::CC_VC;
|
||||
case FEXCore::IR::COND_VS:
|
||||
case FEXCore::IR::COND_VC:
|
||||
case FEXCore::IR::COND_MI: return ARMEmitter::Condition::CC_MI;
|
||||
case FEXCore::IR::COND_PL: return ARMEmitter::Condition::CC_PL;
|
||||
default: LOGMAN_MSG_A_FMT("Unsupported compare type"); return ARMEmitter::Condition::CC_NV;
|
||||
}
|
||||
}
|
||||
|
||||
DEF_OP(CondJump) {
|
||||
auto Op = IROp->C<IR::IROp_CondJump>();
|
||||
|
||||
auto TrueTargetLabel = &JumpTargets.try_emplace(Op->TrueBlock.ID()).first->second;
|
||||
|
||||
if (Op->FromNZCV) {
|
||||
b(MapBranchCC(Op->Cond), TrueTargetLabel);
|
||||
b(MapCC(Op->Cond), TrueTargetLabel);
|
||||
} else {
|
||||
[[maybe_unused]] uint64_t Const;
|
||||
[[maybe_unused]] const bool isConst = IsInlineConstant(Op->Cmp2, &Const);
|
||||
|
||||
@@ -5,7 +5,6 @@ tags: backend|arm64
|
||||
$end_info$
|
||||
*/
|
||||
|
||||
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
|
||||
#include "Interface/Core/JIT/Arm64/JITClass.h"
|
||||
|
||||
namespace FEXCore::CPU {
|
||||
@@ -18,12 +17,7 @@ DEF_OP(VInsGPR) {
|
||||
const auto ElementSize = Op->Header.ElementSize;
|
||||
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(ElementSize == 8 || ElementSize == 4 || ElementSize == 2 || ElementSize == 1, "Unexpected {} size", __func__);
|
||||
const auto SubEmitSize = ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
|
||||
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
|
||||
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
|
||||
ElementSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
|
||||
ARMEmitter::SubRegSize::i8Bit;
|
||||
const auto SubEmitSize = ConvertSubRegSize8(IROp);
|
||||
const auto ElementsPer128Bit = 16 / ElementSize;
|
||||
|
||||
const auto Dst = GetVReg(Node);
|
||||
@@ -66,7 +60,7 @@ DEF_OP(VInsGPR) {
|
||||
// Inserts the GPR value into the given V register.
|
||||
// Also automatically adjusts the index in the case of using the
|
||||
// moved upper lane.
|
||||
const auto Insert = [&](const FEXCore::ARMEmitter::VRegister& reg, int index) {
|
||||
const auto Insert = [&](const ARMEmitter::VRegister& reg, int index) {
|
||||
if (InUpperLane) {
|
||||
index -= ElementsPer128Bit;
|
||||
}
|
||||
@@ -117,16 +111,7 @@ DEF_OP(VDupFromGPR) {
|
||||
const auto Src = GetReg(Op->Src.ID());
|
||||
|
||||
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
|
||||
const auto ElementSize = IROp->ElementSize;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(ElementSize == 8 || ElementSize == 4 || ElementSize == 2 || ElementSize == 1, "Unexpected {} element size: {}",
|
||||
__func__, ElementSize);
|
||||
|
||||
const auto SubEmitSize = ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
|
||||
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
|
||||
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
|
||||
ElementSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
|
||||
ARMEmitter::SubRegSize::i8Bit;
|
||||
const auto SubEmitSize = ConvertSubRegSize8(IROp);
|
||||
|
||||
if (HostSupportsSVE256 && Is256Bit) {
|
||||
dup(SubEmitSize, Dst.Z(), Src);
|
||||
@@ -216,14 +201,9 @@ DEF_OP(Vector_SToF) {
|
||||
const auto OpSize = IROp->Size;
|
||||
|
||||
const auto ElementSize = Op->Header.ElementSize;
|
||||
const auto SubEmitSize = ConvertSubRegSize248(IROp);
|
||||
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(ElementSize == 8 || ElementSize == 4 || ElementSize == 2, "Unexpected {} size", __func__);
|
||||
const auto SubEmitSize = ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
|
||||
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
|
||||
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
|
||||
ARMEmitter::SubRegSize::i16Bit;
|
||||
|
||||
const auto Dst = GetVReg(Node);
|
||||
const auto Vector = GetVReg(Op->Vector.ID());
|
||||
if (HostSupportsSVE256 && Is256Bit) {
|
||||
@@ -253,14 +233,9 @@ DEF_OP(Vector_FToZS) {
|
||||
const auto OpSize = IROp->Size;
|
||||
|
||||
const auto ElementSize = Op->Header.ElementSize;
|
||||
const auto SubEmitSize = ConvertSubRegSize248(IROp);
|
||||
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(ElementSize == 8 || ElementSize == 4 || ElementSize == 2, "Unexpected {} size", __func__);
|
||||
const auto SubEmitSize = ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
|
||||
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
|
||||
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
|
||||
ARMEmitter::SubRegSize::i16Bit;
|
||||
|
||||
const auto Dst = GetVReg(Node);
|
||||
const auto Vector = GetVReg(Op->Vector.ID());
|
||||
if (HostSupportsSVE256 && Is256Bit) {
|
||||
@@ -289,14 +264,8 @@ DEF_OP(Vector_FToS) {
|
||||
const auto Op = IROp->C<IR::IROp_Vector_FToS>();
|
||||
const auto OpSize = IROp->Size;
|
||||
|
||||
const auto ElementSize = Op->Header.ElementSize;
|
||||
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(ElementSize == 8 || ElementSize == 4 || ElementSize == 2, "Unexpected {} size", __func__);
|
||||
const auto SubEmitSize = ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
|
||||
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
|
||||
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
|
||||
ARMEmitter::SubRegSize::i16Bit;
|
||||
const auto SubEmitSize = ConvertSubRegSize248(IROp);
|
||||
|
||||
const auto Dst = GetVReg(Node);
|
||||
const auto Vector = GetVReg(Op->Vector.ID());
|
||||
@@ -323,15 +292,10 @@ DEF_OP(Vector_FToF) {
|
||||
const auto OpSize = IROp->Size;
|
||||
|
||||
const auto ElementSize = Op->Header.ElementSize;
|
||||
const auto SubEmitSize = ConvertSubRegSize248(IROp);
|
||||
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
|
||||
const auto Conv = (ElementSize << 8) | Op->SrcElementSize;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(ElementSize == 8 || ElementSize == 4 || ElementSize == 2, "Unexpected {} size", __func__);
|
||||
const auto SubEmitSize = ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
|
||||
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
|
||||
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
|
||||
ARMEmitter::SubRegSize::i16Bit;
|
||||
|
||||
const auto Dst = GetVReg(Node);
|
||||
const auto Vector = GetVReg(Op->Vector.ID());
|
||||
|
||||
@@ -353,23 +317,23 @@ DEF_OP(Vector_FToF) {
|
||||
|
||||
switch (Conv) {
|
||||
case 0x0402: { // Float <- Half
|
||||
zip1(FEXCore::ARMEmitter::SubRegSize::i16Bit, Dst.Z(), Vector.Z(), Vector.Z());
|
||||
fcvtlt(FEXCore::ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Mask, Dst.Z());
|
||||
zip1(ARMEmitter::SubRegSize::i16Bit, Dst.Z(), Vector.Z(), Vector.Z());
|
||||
fcvtlt(ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Mask, Dst.Z());
|
||||
break;
|
||||
}
|
||||
case 0x0804: { // Double <- Float
|
||||
zip1(FEXCore::ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Vector.Z(), Vector.Z());
|
||||
fcvtlt(FEXCore::ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Dst.Z());
|
||||
zip1(ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Vector.Z(), Vector.Z());
|
||||
fcvtlt(ARMEmitter::SubRegSize::i64Bit, Dst.Z(), Mask, Dst.Z());
|
||||
break;
|
||||
}
|
||||
case 0x0204: { // Half <- Float
|
||||
fcvtnt(FEXCore::ARMEmitter::SubRegSize::i16Bit, Dst.Z(), Mask, Vector.Z());
|
||||
uzp2(FEXCore::ARMEmitter::SubRegSize::i16Bit, Dst.Z(), Dst.Z(), Dst.Z());
|
||||
fcvtnt(ARMEmitter::SubRegSize::i16Bit, Dst.Z(), Mask, Vector.Z());
|
||||
uzp2(ARMEmitter::SubRegSize::i16Bit, Dst.Z(), Dst.Z(), Dst.Z());
|
||||
break;
|
||||
}
|
||||
case 0x0408: { // Float <- Double
|
||||
fcvtnt(FEXCore::ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Mask, Vector.Z());
|
||||
uzp2(FEXCore::ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Dst.Z(), Dst.Z());
|
||||
fcvtnt(ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Mask, Vector.Z());
|
||||
uzp2(ARMEmitter::SubRegSize::i32Bit, Dst.Z(), Dst.Z(), Dst.Z());
|
||||
break;
|
||||
}
|
||||
default: LOGMAN_MSG_A_FMT("Unknown Vector_FToF Type : 0x{:04x}", Conv); break;
|
||||
@@ -391,18 +355,46 @@ DEF_OP(Vector_FToF) {
|
||||
}
|
||||
}
|
||||
|
||||
DEF_OP(VFCVTL2) {
|
||||
const auto Op = IROp->C<IR::IROp_VFCVTL2>();
|
||||
|
||||
const auto SubEmitSize = ConvertSubRegSize248(IROp);
|
||||
|
||||
const auto Dst = GetVReg(Node);
|
||||
const auto Vector = GetVReg(Op->Vector.ID());
|
||||
|
||||
fcvtl2(SubEmitSize, Dst.D(), Vector.D());
|
||||
}
|
||||
|
||||
DEF_OP(VFCVTN2) {
|
||||
const auto Op = IROp->C<IR::IROp_VFCVTN2>();
|
||||
|
||||
const auto SubEmitSize = ConvertSubRegSize248(IROp);
|
||||
|
||||
const auto Dst = GetVReg(Node);
|
||||
const auto VectorLower = GetVReg(Op->VectorLower.ID());
|
||||
const auto VectorUpper = GetVReg(Op->VectorUpper.ID());
|
||||
|
||||
auto Lower = VectorLower;
|
||||
if (Dst != VectorLower) {
|
||||
mov(VTMP1.Q(), VectorLower.Q());
|
||||
Lower = VTMP1;
|
||||
}
|
||||
|
||||
fcvtn2(SubEmitSize, Lower.Q(), VectorUpper.Q());
|
||||
|
||||
if (Dst != VectorLower) {
|
||||
mov(Dst.Q(), Lower.Q());
|
||||
}
|
||||
}
|
||||
|
||||
DEF_OP(Vector_FToI) {
|
||||
const auto Op = IROp->C<IR::IROp_Vector_FToI>();
|
||||
const auto OpSize = IROp->Size;
|
||||
|
||||
const auto ElementSize = Op->Header.ElementSize;
|
||||
const auto SubEmitSize = ConvertSubRegSize248(IROp);
|
||||
const auto Is256Bit = OpSize == Core::CPUState::XMM_AVX_REG_SIZE;
|
||||
LOGMAN_THROW_AA_FMT(ElementSize == 8 || ElementSize == 4 || ElementSize == 2, "Unexpected {} size", __func__);
|
||||
|
||||
const auto SubEmitSize = ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
|
||||
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
|
||||
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
|
||||
ARMEmitter::SubRegSize::i16Bit;
|
||||
|
||||
const auto Dst = GetVReg(Node);
|
||||
const auto Vector = GetVReg(Op->Vector.ID());
|
||||
@@ -425,15 +417,15 @@ DEF_OP(Vector_FToI) {
|
||||
// frinti having AdvSIMD, AdvSIMD scalar, and an SVE version),
|
||||
// we can't just use a lambda without some seriously ugly casting.
|
||||
// This is fairly self-contained otherwise.
|
||||
#define ROUNDING_FN(name) \
|
||||
if (ElementSize == 2) { \
|
||||
name(Dst.H(), Vector.H()); \
|
||||
#define ROUNDING_FN(name) \
|
||||
if (ElementSize == 2) { \
|
||||
name(Dst.H(), Vector.H()); \
|
||||
} else if (ElementSize == 4) { \
|
||||
name(Dst.S(), Vector.S()); \
|
||||
name(Dst.S(), Vector.S()); \
|
||||
} else if (ElementSize == 8) { \
|
||||
name(Dst.D(), Vector.D()); \
|
||||
} else { \
|
||||
FEX_UNREACHABLE; \
|
||||
name(Dst.D(), Vector.D()); \
|
||||
} else { \
|
||||
FEX_UNREACHABLE; \
|
||||
}
|
||||
|
||||
switch (Op->Round) {
|
||||
|
||||
@@ -5,9 +5,7 @@ tags: backend|arm64
|
||||
$end_info$
|
||||
*/
|
||||
|
||||
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
|
||||
#include "Interface/Core/JIT/Arm64/JITClass.h"
|
||||
#include "Interface/IR/Passes/RegisterAllocationPass.h"
|
||||
|
||||
namespace FEXCore::CPU {
|
||||
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const* IROp, IR::NodeID Node)
|
||||
|
||||
@@ -13,7 +13,6 @@ $end_info$
|
||||
|
||||
#include "FEXCore/Utils/Telemetry.h"
|
||||
#include "Interface/Context/Context.h"
|
||||
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
|
||||
#include "Interface/Core/LookupCache.h"
|
||||
|
||||
#include "Interface/Core/Dispatcher/Dispatcher.h"
|
||||
@@ -469,13 +468,13 @@ void Arm64JITCore::Op_Unhandled(const IR::IROp_Header* IROp, IR::NodeID Node) {
|
||||
static void DirectBlockDelinker(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
|
||||
auto LinkerAddress = Frame->Pointers.Common.ExitFunctionLinker;
|
||||
uintptr_t branch = (uintptr_t)(Record)-8;
|
||||
FEXCore::ARMEmitter::Emitter emit((uint8_t*)(branch), 8);
|
||||
FEXCore::ARMEmitter::SingleUseForwardLabel l_BranchHost;
|
||||
ARMEmitter::Emitter emit((uint8_t*)(branch), 8);
|
||||
ARMEmitter::SingleUseForwardLabel l_BranchHost;
|
||||
emit.ldr(TMP1, &l_BranchHost);
|
||||
emit.blr(TMP1);
|
||||
emit.Bind(&l_BranchHost);
|
||||
emit.dc64(LinkerAddress);
|
||||
FEXCore::ARMEmitter::Emitter::ClearICache((void*)branch, 8);
|
||||
ARMEmitter::Emitter::ClearICache((void*)branch, 8);
|
||||
}
|
||||
|
||||
static void IndirectBlockDelinker(FEXCore::Core::CpuStateFrame* Frame, FEXCore::Context::ExitFunctionLinkData* Record) {
|
||||
@@ -500,9 +499,9 @@ static uint64_t Arm64JITCore_ExitFunctionLink(FEXCore::Core::CpuStateFrame* Fram
|
||||
if (vixl::IsInt26(offset)) {
|
||||
// optimal case - can branch directly
|
||||
// patch the code
|
||||
FEXCore::ARMEmitter::Emitter emit((uint8_t*)(branch), 4);
|
||||
ARMEmitter::Emitter emit((uint8_t*)(branch), 4);
|
||||
emit.b(offset);
|
||||
FEXCore::ARMEmitter::Emitter::ClearICache((void*)branch, 4);
|
||||
ARMEmitter::Emitter::ClearICache((void*)branch, 4);
|
||||
|
||||
// Add de-linking handler
|
||||
Thread->LookupCache->AddBlockLink(GuestRip, Record, DirectBlockDelinker);
|
||||
@@ -522,27 +521,20 @@ void Arm64JITCore::Op_NoOp(const IR::IROp_Header* IROp, IR::NodeID Node) {}
|
||||
Arm64JITCore::Arm64JITCore(FEXCore::Context::ContextImpl* ctx, FEXCore::Core::InternalThreadState* Thread)
|
||||
: CPUBackend(Thread, INITIAL_CODE_SIZE, MAX_CODE_SIZE)
|
||||
, Arm64Emitter(ctx)
|
||||
, HostSupportsSVE128 {ctx->HostFeatures.SupportsSVE}
|
||||
, HostSupportsSVE256 {ctx->HostFeatures.SupportsAVX}
|
||||
, HostSupportsSVE128 {ctx->HostFeatures.SupportsSVE128}
|
||||
, HostSupportsSVE256 {ctx->HostFeatures.SupportsSVE256}
|
||||
, HostSupportsAVX256 {ctx->HostFeatures.SupportsAVX && ctx->HostFeatures.SupportsSVE256}
|
||||
, HostSupportsRPRES {ctx->HostFeatures.SupportsRPRES}
|
||||
, HostSupportsAFP {ctx->HostFeatures.SupportsAFP}
|
||||
, CTX {ctx} {
|
||||
|
||||
RAPass = Thread->PassManager->GetPass<IR::RegisterAllocationPass>("RA");
|
||||
|
||||
RAPass->AllocateRegisterSet(RegisterClasses);
|
||||
|
||||
RAPass->AddRegisters(FEXCore::IR::GPRClass, GeneralRegisters.size());
|
||||
RAPass->AddRegisters(FEXCore::IR::GPRFixedClass, StaticRegisters.size());
|
||||
RAPass->AddRegisters(FEXCore::IR::FPRClass, GeneralFPRegisters.size());
|
||||
RAPass->AddRegisters(FEXCore::IR::FPRFixedClass, StaticFPRegisters.size());
|
||||
RAPass->AddRegisters(FEXCore::IR::GPRPairClass, GeneralPairRegisters.size());
|
||||
RAPass->AddRegisters(FEXCore::IR::ComplexClass, 1);
|
||||
|
||||
for (uint32_t i = 0; i < GeneralPairRegisters.size(); ++i) {
|
||||
RAPass->AddRegisterConflict(FEXCore::IR::GPRClass, i * 2, FEXCore::IR::GPRPairClass, i);
|
||||
RAPass->AddRegisterConflict(FEXCore::IR::GPRClass, i * 2 + 1, FEXCore::IR::GPRPairClass, i);
|
||||
}
|
||||
RAPass->PairRegs = PairRegisters;
|
||||
|
||||
{
|
||||
// Set up pointers that the JIT needs to load
|
||||
@@ -669,7 +661,7 @@ bool Arm64JITCore::IsGPRPair(IR::NodeID Node) const {
|
||||
}
|
||||
|
||||
CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, const FEXCore::IR::IRListView* IR, FEXCore::Core::DebugData* DebugData,
|
||||
FEXCore::IR::RegisterAllocationData* RAData) {
|
||||
const FEXCore::IR::RegisterAllocationData* RAData) {
|
||||
FEXCORE_PROFILE_SCOPED("Arm64::CompileCode");
|
||||
|
||||
JumpTargets.clear();
|
||||
@@ -732,8 +724,6 @@ CPUBackend::CompiledCode Arm64JITCore::CompileCode(uint64_t Entry, const FEXCore
|
||||
offsetof(FEXCore::Core::InternalThreadState, InterruptFaultPage) - offsetof(FEXCore::Core::InternalThreadState, BaseFrameState));
|
||||
}
|
||||
|
||||
// LOGMAN_THROW_A_FMT(RAData->HasFullRA(), "Arm64 JIT only works with RA");
|
||||
|
||||
SpillSlots = RAData->SpillSlots();
|
||||
|
||||
if (SpillSlots) {
|
||||
@@ -898,7 +888,6 @@ fextl::unique_ptr<CPUBackend> CreateArm64JITCore(FEXCore::Context::ContextImpl*
|
||||
CPUBackendFeatures GetArm64JITBackendFeatures() {
|
||||
return CPUBackendFeatures {
|
||||
.SupportsFlags = true,
|
||||
.SupportsSaturatingRoundingShifts = true,
|
||||
.SupportsVTBL2 = true,
|
||||
};
|
||||
}
|
||||
|
||||
@@ -8,7 +8,6 @@ $end_info$
|
||||
#pragma once
|
||||
|
||||
#include "Interface/Core/ArchHelpers/Arm64Emitter.h"
|
||||
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
|
||||
#include "Interface/Core/CPUBackend.h"
|
||||
#include "Interface/Core/Dispatcher/Dispatcher.h"
|
||||
#include "Interface/IR/IR.h"
|
||||
@@ -24,6 +23,8 @@ $end_info$
|
||||
#include <FEXCore/fextl/string.h>
|
||||
#include <FEXCore/fextl/vector.h>
|
||||
|
||||
#include <CodeEmitter/Emitter.h>
|
||||
|
||||
#include <array>
|
||||
#include <cstdint>
|
||||
#include <utility>
|
||||
@@ -46,7 +47,7 @@ public:
|
||||
|
||||
[[nodiscard]]
|
||||
CPUBackend::CompiledCode CompileCode(uint64_t Entry, const FEXCore::IR::IRListView* IR, FEXCore::Core::DebugData* DebugData,
|
||||
FEXCore::IR::RegisterAllocationData* RAData) override;
|
||||
const FEXCore::IR::RegisterAllocationData* RAData) override;
|
||||
|
||||
[[nodiscard]]
|
||||
void* MapRegion(void* HostPtr, uint64_t, uint64_t) override {
|
||||
@@ -71,6 +72,7 @@ private:
|
||||
|
||||
const bool HostSupportsSVE128 {};
|
||||
const bool HostSupportsSVE256 {};
|
||||
const bool HostSupportsAVX256 {};
|
||||
const bool HostSupportsRPRES {};
|
||||
const bool HostSupportsAFP {};
|
||||
|
||||
@@ -83,7 +85,7 @@ private:
|
||||
fextl::map<IR::NodeID, ARMEmitter::BiDirectionalLabel> JumpTargets;
|
||||
|
||||
[[nodiscard]]
|
||||
FEXCore::ARMEmitter::Register GetReg(IR::NodeID Node) const {
|
||||
ARMEmitter::Register GetReg(IR::NodeID Node) const {
|
||||
const auto Reg = GetPhys(Node);
|
||||
|
||||
LOGMAN_THROW_AA_FMT(Reg.Class == IR::GPRFixedClass.Val || Reg.Class == IR::GPRClass.Val, "Unexpected Class: {}", Reg.Class);
|
||||
@@ -98,7 +100,7 @@ private:
|
||||
}
|
||||
|
||||
[[nodiscard]]
|
||||
FEXCore::ARMEmitter::VRegister GetVReg(IR::NodeID Node) const {
|
||||
ARMEmitter::VRegister GetVReg(IR::NodeID Node) const {
|
||||
const auto Reg = GetPhys(Node);
|
||||
|
||||
LOGMAN_THROW_AA_FMT(Reg.Class == IR::FPRFixedClass.Val || Reg.Class == IR::FPRClass.Val, "Unexpected Class: {}", Reg.Class);
|
||||
@@ -113,12 +115,12 @@ private:
|
||||
}
|
||||
|
||||
[[nodiscard]]
|
||||
std::pair<FEXCore::ARMEmitter::Register, FEXCore::ARMEmitter::Register> GetRegPair(IR::NodeID Node) const {
|
||||
std::pair<ARMEmitter::Register, ARMEmitter::Register> GetRegPair(IR::NodeID Node) const {
|
||||
const auto Reg = GetPhys(Node);
|
||||
|
||||
LOGMAN_THROW_AA_FMT(Reg.Class == IR::GPRPairClass.Val, "Unexpected Class: {}", Reg.Class);
|
||||
|
||||
return GeneralPairRegisters[Reg.Reg];
|
||||
return std::make_pair(GeneralRegisters[Reg.Reg], GeneralRegisters[Reg.Reg + 1]);
|
||||
}
|
||||
|
||||
[[nodiscard]]
|
||||
@@ -134,7 +136,7 @@ private:
|
||||
}
|
||||
|
||||
[[nodiscard]]
|
||||
FEXCore::ARMEmitter::Register GetZeroableReg(IR::OrderedNodeWrapper Src) const {
|
||||
ARMEmitter::Register GetZeroableReg(IR::OrderedNodeWrapper Src) const {
|
||||
uint64_t Const;
|
||||
if (IsInlineConstant(Src, &Const)) {
|
||||
LOGMAN_THROW_AA_FMT(Const == 0, "Only valid constant");
|
||||
@@ -154,6 +156,99 @@ private:
|
||||
ARMEmitter::ShiftType::ROR;
|
||||
}
|
||||
|
||||
[[nodiscard]]
|
||||
ARMEmitter::Size ConvertSize(const IR::IROp_Header* Op) {
|
||||
return Op->Size == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
}
|
||||
|
||||
[[nodiscard]]
|
||||
ARMEmitter::Size ConvertSize48(const IR::IROp_Header* Op) {
|
||||
LOGMAN_THROW_AA_FMT(Op->Size == 4 || Op->Size == 8, "Invalid size");
|
||||
return ConvertSize(Op);
|
||||
}
|
||||
|
||||
[[nodiscard]]
|
||||
ARMEmitter::SubRegSize ConvertSubRegSize16(uint8_t ElementSize) {
|
||||
LOGMAN_THROW_AA_FMT(ElementSize == 1 || ElementSize == 2 || ElementSize == 4 || ElementSize == 8 || ElementSize == 16, "Invalid size");
|
||||
return ElementSize == 1 ? ARMEmitter::SubRegSize::i8Bit :
|
||||
ElementSize == 2 ? ARMEmitter::SubRegSize::i16Bit :
|
||||
ElementSize == 4 ? ARMEmitter::SubRegSize::i32Bit :
|
||||
ElementSize == 8 ? ARMEmitter::SubRegSize::i64Bit :
|
||||
ARMEmitter::SubRegSize::i128Bit;
|
||||
}
|
||||
|
||||
[[nodiscard]]
|
||||
ARMEmitter::SubRegSize ConvertSubRegSize16(const IR::IROp_Header* Op) {
|
||||
return ConvertSubRegSize16(Op->ElementSize);
|
||||
}
|
||||
|
||||
[[nodiscard]]
|
||||
ARMEmitter::SubRegSize ConvertSubRegSize8(uint8_t ElementSize) {
|
||||
LOGMAN_THROW_AA_FMT(ElementSize != 16, "Invalid size");
|
||||
return ConvertSubRegSize16(ElementSize);
|
||||
}
|
||||
|
||||
[[nodiscard]]
|
||||
ARMEmitter::SubRegSize ConvertSubRegSize8(const IR::IROp_Header* Op) {
|
||||
return ConvertSubRegSize8(Op->ElementSize);
|
||||
}
|
||||
|
||||
[[nodiscard]]
|
||||
ARMEmitter::SubRegSize ConvertSubRegSize4(const IR::IROp_Header* Op) {
|
||||
LOGMAN_THROW_AA_FMT(Op->ElementSize != 8, "Invalid size");
|
||||
return ConvertSubRegSize8(Op);
|
||||
}
|
||||
|
||||
[[nodiscard]]
|
||||
ARMEmitter::SubRegSize ConvertSubRegSize248(const IR::IROp_Header* Op) {
|
||||
LOGMAN_THROW_AA_FMT(Op->ElementSize != 1, "Invalid size");
|
||||
return ConvertSubRegSize8(Op);
|
||||
}
|
||||
|
||||
[[nodiscard]]
|
||||
ARMEmitter::VectorRegSizePair ConvertSubRegSizePair16(const IR::IROp_Header* Op) {
|
||||
return ARMEmitter::ToVectorSizePair(ConvertSubRegSize16(Op));
|
||||
}
|
||||
|
||||
[[nodiscard]]
|
||||
ARMEmitter::VectorRegSizePair ConvertSubRegSizePair8(const IR::IROp_Header* Op) {
|
||||
LOGMAN_THROW_AA_FMT(Op->ElementSize != 16, "Invalid size");
|
||||
return ConvertSubRegSizePair16(Op);
|
||||
}
|
||||
|
||||
[[nodiscard]]
|
||||
ARMEmitter::VectorRegSizePair ConvertSubRegSizePair248(const IR::IROp_Header* Op) {
|
||||
LOGMAN_THROW_AA_FMT(Op->ElementSize != 1, "Invalid size");
|
||||
return ConvertSubRegSizePair8(Op);
|
||||
}
|
||||
|
||||
[[nodiscard]]
|
||||
ARMEmitter::Condition MapCC(IR::CondClassType Cond) {
|
||||
switch (Cond.Val) {
|
||||
case FEXCore::IR::COND_EQ: return ARMEmitter::Condition::CC_EQ;
|
||||
case FEXCore::IR::COND_NEQ: return ARMEmitter::Condition::CC_NE;
|
||||
case FEXCore::IR::COND_SGE: return ARMEmitter::Condition::CC_GE;
|
||||
case FEXCore::IR::COND_SLT: return ARMEmitter::Condition::CC_LT;
|
||||
case FEXCore::IR::COND_SGT: return ARMEmitter::Condition::CC_GT;
|
||||
case FEXCore::IR::COND_SLE: return ARMEmitter::Condition::CC_LE;
|
||||
case FEXCore::IR::COND_UGE: return ARMEmitter::Condition::CC_CS;
|
||||
case FEXCore::IR::COND_ULT: return ARMEmitter::Condition::CC_CC;
|
||||
case FEXCore::IR::COND_UGT: return ARMEmitter::Condition::CC_HI;
|
||||
case FEXCore::IR::COND_ULE: return ARMEmitter::Condition::CC_LS;
|
||||
case FEXCore::IR::COND_FLU: return ARMEmitter::Condition::CC_LT;
|
||||
case FEXCore::IR::COND_FGE: return ARMEmitter::Condition::CC_GE;
|
||||
case FEXCore::IR::COND_FLEU: return ARMEmitter::Condition::CC_LE;
|
||||
case FEXCore::IR::COND_FGT: return ARMEmitter::Condition::CC_GT;
|
||||
case FEXCore::IR::COND_FU: return ARMEmitter::Condition::CC_VS;
|
||||
case FEXCore::IR::COND_FNU: return ARMEmitter::Condition::CC_VC;
|
||||
case FEXCore::IR::COND_VS:
|
||||
case FEXCore::IR::COND_VC:
|
||||
case FEXCore::IR::COND_MI: return ARMEmitter::Condition::CC_MI;
|
||||
case FEXCore::IR::COND_PL: return ARMEmitter::Condition::CC_PL;
|
||||
default: LOGMAN_MSG_A_FMT("Unsupported compare type"); return ARMEmitter::Condition::CC_NV;
|
||||
}
|
||||
}
|
||||
|
||||
[[nodiscard]]
|
||||
bool IsFPR(IR::NodeID Node) const;
|
||||
[[nodiscard]]
|
||||
@@ -162,8 +257,8 @@ private:
|
||||
bool IsGPRPair(IR::NodeID Node) const;
|
||||
|
||||
[[nodiscard]]
|
||||
FEXCore::ARMEmitter::ExtendedMemOperand GenerateMemOperand(
|
||||
uint8_t AccessSize, FEXCore::ARMEmitter::Register Base, IR::OrderedNodeWrapper Offset, IR::MemOffsetType OffsetType, uint8_t OffsetScale);
|
||||
ARMEmitter::ExtendedMemOperand GenerateMemOperand(uint8_t AccessSize, ARMEmitter::Register Base, IR::OrderedNodeWrapper Offset,
|
||||
IR::MemOffsetType OffsetType, uint8_t OffsetScale);
|
||||
|
||||
// NOTE: Will use TMP1 as a way to encode immediates that happen to fall outside
|
||||
// the limits of the scalar plus immediate variant of SVE load/stores.
|
||||
@@ -171,8 +266,8 @@ private:
|
||||
// TMP1 is safe to use again once this memory operand is used with its
|
||||
// equivalent loads or stores that this was called for.
|
||||
[[nodiscard]]
|
||||
FEXCore::ARMEmitter::SVEMemOperand GenerateSVEMemOperand(uint8_t AccessSize, FEXCore::ARMEmitter::Register Base,
|
||||
IR::OrderedNodeWrapper Offset, IR::MemOffsetType OffsetType, uint8_t OffsetScale);
|
||||
ARMEmitter::SVEMemOperand GenerateSVEMemOperand(uint8_t AccessSize, ARMEmitter::Register Base, IR::OrderedNodeWrapper Offset,
|
||||
IR::MemOffsetType OffsetType, uint8_t OffsetScale);
|
||||
|
||||
[[nodiscard]]
|
||||
bool IsInlineConstant(const IR::OrderedNodeWrapper& Node, uint64_t* Value = nullptr) const;
|
||||
@@ -187,7 +282,7 @@ private:
|
||||
// This is purely a debugging aid for developers to see if they are in JIT code space when inspecting raw memory
|
||||
void EmitDetectionString();
|
||||
IR::RegisterAllocationPass* RAPass;
|
||||
IR::RegisterAllocationData* RAData;
|
||||
const IR::RegisterAllocationData* RAData;
|
||||
FEXCore::Core::DebugData* DebugData;
|
||||
|
||||
void ResetStack();
|
||||
|
||||
File diff suppressed because it is too large.
Load diff
@@ -10,7 +10,6 @@ $end_info$
|
||||
#endif
|
||||
|
||||
#include "Interface/Context/Context.h"
|
||||
#include "Interface/Core/ArchHelpers/CodeEmitter/Emitter.h"
|
||||
#include "Interface/Core/JIT/Arm64/JITClass.h"
|
||||
#include "FEXCore/Debug/InternalThreadState.h"
|
||||
|
||||
@@ -28,9 +27,9 @@ DEF_OP(GuestOpcode) {
|
||||
DEF_OP(Fence) {
|
||||
auto Op = IROp->C<IR::IROp_Fence>();
|
||||
switch (Op->Fence) {
|
||||
case IR::Fence_Load.Val: dmb(FEXCore::ARMEmitter::BarrierScope::LD); break;
|
||||
case IR::Fence_LoadStore.Val: dmb(FEXCore::ARMEmitter::BarrierScope::SY); break;
|
||||
case IR::Fence_Store.Val: dmb(FEXCore::ARMEmitter::BarrierScope::ST); break;
|
||||
case IR::Fence_Load.Val: dmb(ARMEmitter::BarrierScope::LD); break;
|
||||
case IR::Fence_LoadStore.Val: dmb(ARMEmitter::BarrierScope::SY); break;
|
||||
case IR::Fence_Store.Val: dmb(ARMEmitter::BarrierScope::ST); break;
|
||||
default: LOGMAN_MSG_A_FMT("Unknown Fence: {}", Op->Fence); break;
|
||||
}
|
||||
}
|
||||
@@ -121,6 +120,39 @@ DEF_OP(SetRoundingMode) {
|
||||
msr(ARMEmitter::SystemRegister::FPCR, TMP1);
|
||||
}
|
||||
|
||||
DEF_OP(PushRoundingMode) {
|
||||
auto Op = IROp->C<IR::IROp_PushRoundingMode>();
|
||||
auto Dest = GetReg(Node);
|
||||
|
||||
// Save the old rounding mode
|
||||
mrs(Dest, ARMEmitter::SystemRegister::FPCR);
|
||||
|
||||
// vixl simulator doesn't support anything beyond ties-to-even rounding
|
||||
if (!CTX->Config.DisableVixlIndirectCalls) [[unlikely]] {
|
||||
return;
|
||||
}
|
||||
|
||||
// Insert the rounding flags, reversing the mode bits as above
|
||||
if (Op->RoundMode == 3) {
|
||||
orr(ARMEmitter::Size::i64Bit, TMP1, Dest, 3 << 22);
|
||||
} else if (Op->RoundMode == 0) {
|
||||
and_(ARMEmitter::Size::i64Bit, TMP1, Dest, ~(3 << 22));
|
||||
} else {
|
||||
LOGMAN_THROW_AA_FMT(Op->RoundMode == 1 || Op->RoundMode == 2, "expect a valid round mode");
|
||||
|
||||
and_(ARMEmitter::Size::i64Bit, TMP1, Dest, ~(Op->RoundMode << 22));
|
||||
orr(ARMEmitter::Size::i64Bit, TMP1, TMP1, (Op->RoundMode == 2 ? 1 : 2) << 22);
|
||||
}
|
||||
|
||||
// Now save the new FPCR
|
||||
msr(ARMEmitter::SystemRegister::FPCR, TMP1);
|
||||
}
|
||||
|
||||
DEF_OP(PopRoundingMode) {
|
||||
auto Op = IROp->C<IR::IROp_PopRoundingMode>();
|
||||
msr(ARMEmitter::SystemRegister::FPCR, GetReg(Op->FPCR.ID()));
|
||||
}
|
||||
|
||||
DEF_OP(Print) {
|
||||
auto Op = IROp->C<IR::IROp_Print>();
|
||||
|
||||
|
||||
@@ -11,12 +11,14 @@ namespace FEXCore::CPU {
|
||||
#define DEF_OP(x) void Arm64JITCore::Op_##x(IR::IROp_Header const* IROp, IR::NodeID Node)
|
||||
DEF_OP(ExtractElementPair) {
|
||||
auto Op = IROp->C<IR::IROp_ExtractElementPair>();
|
||||
LOGMAN_THROW_AA_FMT(Op->Header.Size == 4 || Op->Header.Size == 8, "Invalid size");
|
||||
const auto EmitSize = Op->Header.Size == 8 ? ARMEmitter::Size::i64Bit : ARMEmitter::Size::i32Bit;
|
||||
|
||||
const auto Src = GetRegPair(Op->Pair.ID());
|
||||
const std::array<ARMEmitter::Register, 2> Regs = {Src.first, Src.second};
|
||||
mov(EmitSize, GetReg(Node), Regs[Op->Element]);
|
||||
const auto Dst = GetReg(Node);
|
||||
const auto Pair = GetRegPair(Op->Pair.ID());
|
||||
const auto Src = Op->Element == 0 ? Pair.first : Pair.second;
|
||||
|
||||
if (Dst != Src) {
|
||||
mov(ConvertSize48(IROp), Dst, Src);
|
||||
}
|
||||
}
|
||||
|
||||
DEF_OP(CreateElementPair) {
|
||||
@@ -42,5 +44,25 @@ DEF_OP(CreateElementPair) {
|
||||
}
|
||||
}
|
||||
|
||||
DEF_OP(Copy) {
|
||||
auto Op = IROp->C<IR::IROp_Copy>();
|
||||
|
||||
mov(ARMEmitter::Size::i64Bit, GetReg(Node), GetReg(Op->Source.ID()));
|
||||
}
|
||||
|
||||
DEF_OP(Swap1) {
|
||||
auto Op = IROp->C<IR::IROp_Swap1>();
|
||||
auto A = GetReg(Op->A.ID()), B = GetReg(Op->B.ID());
|
||||
LOGMAN_THROW_AA_FMT(B == GetReg(Node), "Invariant");
|
||||
|
||||
mov(ARMEmitter::Size::i64Bit, TMP1, A);
|
||||
mov(ARMEmitter::Size::i64Bit, A, B);
|
||||
mov(ARMEmitter::Size::i64Bit, B, TMP1);
|
||||
}
|
||||
|
||||
DEF_OP(Swap2) {
|
||||
// Implemented above
|
||||
}
|
||||
|
||||
#undef DEF_OP
|
||||
} // namespace FEXCore::CPU
|
||||
File diff suppressed because it is too large.
Load diff
File diff suppressed because it is too large.
Load diff
File diff suppressed because it is too large.
Load diff
File diff suppressed because it is too large.
Load diff
@@ -22,10 +22,10 @@ class OrderedNode;
|
||||
#define OpcodeArgs [[maybe_unused]] FEXCore::X86Tables::DecodedOp Op
|
||||
|
||||
void OpDispatchBuilder::SHA1NEXTEOp(OpcodeArgs) {
|
||||
OrderedNode* Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
|
||||
OrderedNode* Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
|
||||
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
|
||||
OrderedNode* RotatedNode {};
|
||||
Ref RotatedNode {};
|
||||
if (CTX->HostFeatures.SupportsSHA) {
|
||||
// ARMv8 SHA1 extension provides a `SHA1H` instruction which does a fixed rotate by 30.
|
||||
// This only operates on element 0 rather than element 3. We don't have the luxury of rewriting the x86 SHA algorithm to take advantage of this.
|
||||
@@ -47,25 +47,25 @@ void OpDispatchBuilder::SHA1NEXTEOp(OpcodeArgs) {
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::SHA1MSG1Op(OpcodeArgs) {
|
||||
OrderedNode* Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
|
||||
OrderedNode* Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
|
||||
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
|
||||
OrderedNode* NewVec = _VExtr(16, 8, Dest, Src, 1);
|
||||
Ref NewVec = _VExtr(16, 8, Dest, Src, 1);
|
||||
|
||||
// [W0, W1, W2, W3] ^ [W2, W3, W4, W5]
|
||||
OrderedNode* Result = _VXor(16, 1, Dest, NewVec);
|
||||
Ref Result = _VXor(16, 1, Dest, NewVec);
|
||||
|
||||
StoreResult(FPRClass, Op, Result, -1);
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::SHA1MSG2Op(OpcodeArgs) {
|
||||
OrderedNode* Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
|
||||
OrderedNode* Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
|
||||
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
|
||||
// This instruction mostly matches ARMv8's SHA1SU1 instruction but one of the elements are flipped in an unexpected way.
|
||||
// Do all the work without it.
|
||||
|
||||
const auto ZeroRegister = LoadAndCacheNamedVectorConstant(OpSize::i32Bit, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_ZERO);
|
||||
const auto ZeroRegister = LoadZeroVector(OpSize::i32Bit);
|
||||
|
||||
// Shift the incoming source left by a 32-bit element, inserting Zeros.
|
||||
// This could be slightly improved to use a VInsGPR with the zero register.
|
||||
@@ -90,20 +90,18 @@ void OpDispatchBuilder::SHA1MSG2Op(OpcodeArgs) {
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
|
||||
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here to indicate function and constants");
|
||||
using FnType = Ref (*)(OpDispatchBuilder&, Ref, Ref, Ref);
|
||||
|
||||
using FnType = OrderedNode* (*)(OpDispatchBuilder&, OrderedNode*, OrderedNode*, OrderedNode*);
|
||||
|
||||
const auto f0 = [](OpDispatchBuilder& Self, OrderedNode* B, OrderedNode* C, OrderedNode* D) -> OrderedNode* {
|
||||
const auto f0 = [](OpDispatchBuilder& Self, Ref B, Ref C, Ref D) -> Ref {
|
||||
return Self._Xor(OpSize::i32Bit, Self._And(OpSize::i32Bit, B, C), Self._Andn(OpSize::i32Bit, D, B));
|
||||
};
|
||||
const auto f1 = [](OpDispatchBuilder& Self, OrderedNode* B, OrderedNode* C, OrderedNode* D) -> OrderedNode* {
|
||||
const auto f1 = [](OpDispatchBuilder& Self, Ref B, Ref C, Ref D) -> Ref {
|
||||
return Self._Xor(OpSize::i32Bit, Self._Xor(OpSize::i32Bit, B, C), D);
|
||||
};
|
||||
const auto f2 = [](OpDispatchBuilder& Self, OrderedNode* B, OrderedNode* C, OrderedNode* D) -> OrderedNode* {
|
||||
const auto f2 = [](OpDispatchBuilder& Self, Ref B, Ref C, Ref D) -> Ref {
|
||||
return Self.BitwiseAtLeastTwo(B, C, D);
|
||||
};
|
||||
const auto f3 = [](OpDispatchBuilder& Self, OrderedNode* B, OrderedNode* C, OrderedNode* D) -> OrderedNode* {
|
||||
const auto f3 = [](OpDispatchBuilder& Self, Ref B, Ref C, Ref D) -> Ref {
|
||||
return Self._Xor(OpSize::i32Bit, Self._Xor(OpSize::i32Bit, B, C), D);
|
||||
};
|
||||
|
||||
@@ -121,16 +119,16 @@ void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
|
||||
f3,
|
||||
};
|
||||
|
||||
const uint64_t Imm8 = Op->Src[1].Data.Literal.Value & 0b11;
|
||||
const uint64_t Imm8 = Op->Src[1].Literal() & 0b11;
|
||||
const FnType Fn = fn_array[Imm8];
|
||||
auto K = _Constant(32, k_array[Imm8]);
|
||||
|
||||
OrderedNode* Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
|
||||
OrderedNode* Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
|
||||
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
|
||||
auto W0E = _VExtractToGPR(16, 4, Src, 3);
|
||||
|
||||
using RoundResult = std::tuple<OrderedNode*, OrderedNode*, OrderedNode*, OrderedNode*, OrderedNode*>;
|
||||
using RoundResult = std::tuple<Ref, Ref, Ref, Ref, Ref>;
|
||||
|
||||
const auto Round0 = [&]() -> RoundResult {
|
||||
auto A = _VExtractToGPR(16, 4, Dest, 3);
|
||||
@@ -147,8 +145,7 @@ void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
|
||||
|
||||
return {A1, B1, C1, D1, E1};
|
||||
};
|
||||
const auto Round1To3 = [&](OrderedNode* A, OrderedNode* B, OrderedNode* C, OrderedNode* D, OrderedNode* E, OrderedNode* Src,
|
||||
unsigned W_idx) -> RoundResult {
|
||||
const auto Round1To3 = [&](Ref A, Ref B, Ref C, Ref D, Ref E, Ref Src, unsigned W_idx) -> RoundResult {
|
||||
// Kill W and E at the beginning
|
||||
auto W = _VExtractToGPR(16, 4, Src, W_idx);
|
||||
auto Q = _Add(OpSize::i32Bit, W, E);
|
||||
@@ -177,15 +174,15 @@ void OpDispatchBuilder::SHA1RNDS4Op(OpcodeArgs) {
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::SHA256MSG1Op(OpcodeArgs) {
|
||||
OrderedNode* Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
|
||||
OrderedNode* Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
|
||||
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
|
||||
OrderedNode* Result {};
|
||||
Ref Result {};
|
||||
|
||||
if (CTX->HostFeatures.SupportsSHA) {
|
||||
Result = _VSha256U0(Dest, Src);
|
||||
} else {
|
||||
const auto Sigma0 = [this](OrderedNode* W) -> OrderedNode* {
|
||||
const auto Sigma0 = [this](Ref W) -> Ref {
|
||||
return _Xor(OpSize::i32Bit, _Xor(OpSize::i32Bit, _Ror(OpSize::i32Bit, W, _Constant(32, 7)), _Ror(OpSize::i32Bit, W, _Constant(32, 18))),
|
||||
_Lshr(OpSize::i32Bit, W, _Constant(32, 3)));
|
||||
};
|
||||
@@ -211,13 +208,13 @@ void OpDispatchBuilder::SHA256MSG1Op(OpcodeArgs) {
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::SHA256MSG2Op(OpcodeArgs) {
|
||||
const auto Sigma1 = [this](OrderedNode* W) -> OrderedNode* {
|
||||
const auto Sigma1 = [this](Ref W) -> Ref {
|
||||
return _Xor(OpSize::i32Bit, _Xor(OpSize::i32Bit, _Ror(OpSize::i32Bit, W, _Constant(32, 17)), _Ror(OpSize::i32Bit, W, _Constant(32, 19))),
|
||||
_Lshr(OpSize::i32Bit, W, _Constant(32, 10)));
|
||||
};
|
||||
|
||||
OrderedNode* Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
|
||||
OrderedNode* Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
|
||||
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
|
||||
auto W14 = _VExtractToGPR(16, 4, Src, 2);
|
||||
auto W15 = _VExtractToGPR(16, 4, Src, 3);
|
||||
@@ -234,7 +231,7 @@ void OpDispatchBuilder::SHA256MSG2Op(OpcodeArgs) {
|
||||
StoreResult(FPRClass, Op, D0, -1);
|
||||
}
|
||||
|
||||
OrderedNode* OpDispatchBuilder::BitwiseAtLeastTwo(OrderedNode* A, OrderedNode* B, OrderedNode* C) {
|
||||
Ref OpDispatchBuilder::BitwiseAtLeastTwo(Ref A, Ref B, Ref C) {
|
||||
// Returns whether at least 2/3 of A/B/C is true.
|
||||
// Expressed as (A & (B | C)) | (B & C)
|
||||
//
|
||||
@@ -245,27 +242,27 @@ OrderedNode* OpDispatchBuilder::BitwiseAtLeastTwo(OrderedNode* A, OrderedNode* B
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::SHA256RNDS2Op(OpcodeArgs) {
|
||||
const auto Ch = [this](OrderedNode* E, OrderedNode* F, OrderedNode* G) -> OrderedNode* {
|
||||
const auto Ch = [this](Ref E, Ref F, Ref G) -> Ref {
|
||||
return _Xor(OpSize::i32Bit, _And(OpSize::i32Bit, E, F), _Andn(OpSize::i32Bit, G, E));
|
||||
};
|
||||
const auto Sigma0 = [this](OrderedNode* A) -> OrderedNode* {
|
||||
const auto Sigma0 = [this](Ref A) -> Ref {
|
||||
return _XorShift(OpSize::i32Bit, _XorShift(OpSize::i32Bit, _Ror(OpSize::i32Bit, A, _Constant(32, 2)), A, ShiftType::ROR, 13), A,
|
||||
ShiftType::ROR, 22);
|
||||
};
|
||||
const auto Sigma1 = [this](OrderedNode* E) -> OrderedNode* {
|
||||
const auto Sigma1 = [this](Ref E) -> Ref {
|
||||
return _XorShift(OpSize::i32Bit, _XorShift(OpSize::i32Bit, _Ror(OpSize::i32Bit, E, _Constant(32, 6)), E, ShiftType::ROR, 11), E,
|
||||
ShiftType::ROR, 25);
|
||||
};
|
||||
|
||||
OrderedNode* Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
|
||||
OrderedNode* Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
|
||||
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
// Hardcoded to XMM0
|
||||
auto XMM0 = LoadXMMRegister(0);
|
||||
|
||||
auto E0 = _VExtractToGPR(16, 4, Src, 1);
|
||||
auto F0 = _VExtractToGPR(16, 4, Src, 0);
|
||||
auto G0 = _VExtractToGPR(16, 4, Dest, 1);
|
||||
OrderedNode* Q0 = _Add(OpSize::i32Bit, Ch(E0, F0, G0), Sigma1(E0));
|
||||
Ref Q0 = _Add(OpSize::i32Bit, Ch(E0, F0, G0), Sigma1(E0));
|
||||
|
||||
auto WK0 = _VExtractToGPR(16, 4, XMM0, 0);
|
||||
Q0 = _Add(OpSize::i32Bit, Q0, WK0);
|
||||
@@ -281,7 +278,7 @@ void OpDispatchBuilder::SHA256RNDS2Op(OpcodeArgs) {
|
||||
auto D0 = _VExtractToGPR(16, 4, Dest, 2);
|
||||
auto E1 = _Add(OpSize::i32Bit, Q0, D0);
|
||||
|
||||
OrderedNode* Q1 = _Add(OpSize::i32Bit, Ch(E1, E0, F0), Sigma1(E1));
|
||||
Ref Q1 = _Add(OpSize::i32Bit, Ch(E1, E0, F0), Sigma1(E1));
|
||||
|
||||
auto WK1 = _VExtractToGPR(16, 4, XMM0, 1);
|
||||
Q1 = _Add(OpSize::i32Bit, Q1, WK1);
|
||||
@@ -305,16 +302,15 @@ void OpDispatchBuilder::SHA256RNDS2Op(OpcodeArgs) {
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::AESImcOp(OpcodeArgs) {
|
||||
OrderedNode* Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
OrderedNode* Result = _VAESImc(Src);
|
||||
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
Ref Result = _VAESImc(Src);
|
||||
StoreResult(FPRClass, Op, Result, -1);
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::AESEncOp(OpcodeArgs) {
|
||||
OrderedNode* Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
|
||||
OrderedNode* Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
const auto ZeroRegister = LoadAndCacheNamedVectorConstant(16, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_ZERO);
|
||||
OrderedNode* Result = _VAESEnc(16, Dest, Src, ZeroRegister);
|
||||
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
|
||||
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
Ref Result = _VAESEnc(16, Dest, Src, LoadZeroVector(16));
|
||||
StoreResult(FPRClass, Op, Result, -1);
|
||||
}
|
||||
|
||||
@@ -325,19 +321,17 @@ void OpDispatchBuilder::VAESEncOp(OpcodeArgs) {
|
||||
// TODO: Handle 256-bit VAESENC.
|
||||
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESENC unimplemented");
|
||||
|
||||
OrderedNode* State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
OrderedNode* Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
|
||||
const auto ZeroRegister = LoadAndCacheNamedVectorConstant(DstSize, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_ZERO);
|
||||
OrderedNode* Result = _VAESEnc(DstSize, State, Key, ZeroRegister);
|
||||
Ref State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
Ref Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
|
||||
Ref Result = _VAESEnc(DstSize, State, Key, LoadZeroVector(DstSize));
|
||||
|
||||
StoreResult(FPRClass, Op, Result, -1);
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::AESEncLastOp(OpcodeArgs) {
|
||||
OrderedNode* Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
|
||||
OrderedNode* Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
const auto ZeroRegister = LoadAndCacheNamedVectorConstant(16, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_ZERO);
|
||||
OrderedNode* Result = _VAESEncLast(16, Dest, Src, ZeroRegister);
|
||||
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
|
||||
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
Ref Result = _VAESEncLast(16, Dest, Src, LoadZeroVector(16));
|
||||
StoreResult(FPRClass, Op, Result, -1);
|
||||
}
|
||||
|
||||
@@ -348,19 +342,17 @@ void OpDispatchBuilder::VAESEncLastOp(OpcodeArgs) {
|
||||
// TODO: Handle 256-bit VAESENCLAST.
|
||||
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESENCLAST unimplemented");
|
||||
|
||||
OrderedNode* State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
OrderedNode* Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
|
||||
const auto ZeroRegister = LoadAndCacheNamedVectorConstant(DstSize, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_ZERO);
|
||||
OrderedNode* Result = _VAESEncLast(DstSize, State, Key, ZeroRegister);
|
||||
Ref State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
Ref Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
|
||||
Ref Result = _VAESEncLast(DstSize, State, Key, LoadZeroVector(DstSize));
|
||||
|
||||
StoreResult(FPRClass, Op, Result, -1);
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::AESDecOp(OpcodeArgs) {
|
||||
OrderedNode* Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
|
||||
OrderedNode* Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
const auto ZeroRegister = LoadAndCacheNamedVectorConstant(16, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_ZERO);
|
||||
OrderedNode* Result = _VAESDec(16, Dest, Src, ZeroRegister);
|
||||
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
|
||||
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
Ref Result = _VAESDec(16, Dest, Src, LoadZeroVector(16));
|
||||
StoreResult(FPRClass, Op, Result, -1);
|
||||
}
|
||||
|
||||
@@ -371,19 +363,17 @@ void OpDispatchBuilder::VAESDecOp(OpcodeArgs) {
|
||||
// TODO: Handle 256-bit VAESDEC.
|
||||
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESDEC unimplemented");
|
||||
|
||||
OrderedNode* State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
OrderedNode* Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
|
||||
const auto ZeroRegister = LoadAndCacheNamedVectorConstant(DstSize, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_ZERO);
|
||||
OrderedNode* Result = _VAESDec(DstSize, State, Key, ZeroRegister);
|
||||
Ref State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
Ref Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
|
||||
Ref Result = _VAESDec(DstSize, State, Key, LoadZeroVector(DstSize));
|
||||
|
||||
StoreResult(FPRClass, Op, Result, -1);
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::AESDecLastOp(OpcodeArgs) {
|
||||
OrderedNode* Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
|
||||
OrderedNode* Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
const auto ZeroRegister = LoadAndCacheNamedVectorConstant(16, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_ZERO);
|
||||
OrderedNode* Result = _VAESDecLast(16, Dest, Src, ZeroRegister);
|
||||
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
|
||||
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
Ref Result = _VAESDecLast(16, Dest, Src, LoadZeroVector(16));
|
||||
StoreResult(FPRClass, Op, Result, -1);
|
||||
}
|
||||
|
||||
@@ -394,50 +384,43 @@ void OpDispatchBuilder::VAESDecLastOp(OpcodeArgs) {
|
||||
// TODO: Handle 256-bit VAESDECLAST.
|
||||
LOGMAN_THROW_A_FMT(Is128Bit, "256-bit VAESDECLAST unimplemented");
|
||||
|
||||
OrderedNode* State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
OrderedNode* Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
|
||||
const auto ZeroRegister = LoadAndCacheNamedVectorConstant(DstSize, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_ZERO);
|
||||
OrderedNode* Result = _VAESDecLast(DstSize, State, Key, ZeroRegister);
|
||||
Ref State = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
Ref Key = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
|
||||
Ref Result = _VAESDecLast(DstSize, State, Key, LoadZeroVector(DstSize));
|
||||
|
||||
StoreResult(FPRClass, Op, Result, -1);
|
||||
}
|
||||
|
||||
OrderedNode* OpDispatchBuilder::AESKeyGenAssistImpl(OpcodeArgs) {
|
||||
OrderedNode* Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Src1 needs to be literal here");
|
||||
const uint64_t RCON = Op->Src[1].Data.Literal.Value;
|
||||
Ref OpDispatchBuilder::AESKeyGenAssistImpl(OpcodeArgs) {
|
||||
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
const uint64_t RCON = Op->Src[1].Literal();
|
||||
|
||||
auto KeyGenSwizzle = LoadAndCacheNamedVectorConstant(16, NAMED_VECTOR_AESKEYGENASSIST_SWIZZLE);
|
||||
const auto ZeroRegister = LoadAndCacheNamedVectorConstant(16, FEXCore::IR::NamedVectorConstant::NAMED_VECTOR_ZERO);
|
||||
return _VAESKeyGenAssist(Src, KeyGenSwizzle, ZeroRegister, RCON);
|
||||
return _VAESKeyGenAssist(Src, KeyGenSwizzle, LoadZeroVector(16), RCON);
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::AESKeyGenAssist(OpcodeArgs) {
|
||||
OrderedNode* Result = AESKeyGenAssistImpl(Op);
|
||||
Ref Result = AESKeyGenAssistImpl(Op);
|
||||
StoreResult(FPRClass, Op, Result, -1);
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::PCLMULQDQOp(OpcodeArgs) {
|
||||
LOGMAN_THROW_A_FMT(Op->Src[1].IsLiteral(), "Selector needs to be literal here");
|
||||
Ref Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
|
||||
Ref Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
const auto Selector = static_cast<uint8_t>(Op->Src[1].Literal());
|
||||
|
||||
OrderedNode* Dest = LoadSource(FPRClass, Op, Op->Dest, Op->Flags);
|
||||
OrderedNode* Src = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
const auto Selector = static_cast<uint8_t>(Op->Src[1].Data.Literal.Value);
|
||||
|
||||
auto Res = _PCLMUL(16, Dest, Src, Selector);
|
||||
auto Res = _PCLMUL(16, Dest, Src, Selector & 0b1'0001);
|
||||
StoreResult(FPRClass, Op, Res, -1);
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::VPCLMULQDQOp(OpcodeArgs) {
|
||||
LOGMAN_THROW_A_FMT(Op->Src[2].IsLiteral(), "Selector needs to be literal here");
|
||||
|
||||
const auto DstSize = GetDstSize(Op);
|
||||
|
||||
OrderedNode* Src1 = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
OrderedNode* Src2 = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
|
||||
const auto Selector = static_cast<uint8_t>(Op->Src[2].Data.Literal.Value);
|
||||
Ref Src1 = LoadSource(FPRClass, Op, Op->Src[0], Op->Flags);
|
||||
Ref Src2 = LoadSource(FPRClass, Op, Op->Src[1], Op->Flags);
|
||||
const auto Selector = static_cast<uint8_t>(Op->Src[2].Literal());
|
||||
|
||||
OrderedNode* Res = _PCLMUL(DstSize, Src1, Src2, Selector);
|
||||
Ref Res = _PCLMUL(DstSize, Src1, Src2, Selector & 0b1'0001);
|
||||
StoreResult(FPRClass, Op, Res, -1);
|
||||
}
|
||||
|
||||
|
||||
@@ -33,7 +33,7 @@ void OpDispatchBuilder::ZeroPF_AF() {
|
||||
SetAF(0);
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::SetPackedRFLAG(bool Lower8, OrderedNode* Src) {
|
||||
void OpDispatchBuilder::SetPackedRFLAG(bool Lower8, Ref Src) {
|
||||
size_t NumFlags = FlagOffsets.size();
|
||||
if (Lower8) {
|
||||
// Calculate flags early.
|
||||
@@ -61,7 +61,7 @@ void OpDispatchBuilder::SetPackedRFLAG(bool Lower8, OrderedNode* Src) {
|
||||
SetRFLAG(Src, FEXCore::X86State::RFLAG_AF_RAW_LOC);
|
||||
} else if (FlagOffset == FEXCore::X86State::RFLAG_PF_RAW_LOC) {
|
||||
// PF is stored parity flipped
|
||||
OrderedNode* Tmp = _Bfe(OpSize::i32Bit, 1, FlagOffset, Src);
|
||||
Ref Tmp = _Bfe(OpSize::i32Bit, 1, FlagOffset, Src);
|
||||
Tmp = _Xor(OpSize::i32Bit, Tmp, _Constant(1));
|
||||
SetRFLAG(Tmp, FlagOffset);
|
||||
} else {
|
||||
@@ -70,11 +70,11 @@ void OpDispatchBuilder::SetPackedRFLAG(bool Lower8, OrderedNode* Src) {
|
||||
}
|
||||
}
|
||||
|
||||
OrderedNode* OpDispatchBuilder::GetPackedRFLAG(uint32_t FlagsMask) {
|
||||
Ref OpDispatchBuilder::GetPackedRFLAG(uint32_t FlagsMask) {
|
||||
// Calculate flags early.
|
||||
CalculateDeferredFlags();
|
||||
|
||||
OrderedNode* Original = _Constant(0);
|
||||
Ref Original = _Constant(0);
|
||||
|
||||
// SF/ZF and N/Z are together on both arm64 and x86_64, so we special case that.
|
||||
bool GetNZ = (FlagsMask & (1 << FEXCore::X86State::RFLAG_SF_RAW_LOC)) && (FlagsMask & (1 << FEXCore::X86State::RFLAG_ZF_RAW_LOC));
|
||||
@@ -99,7 +99,7 @@ OrderedNode* OpDispatchBuilder::GetPackedRFLAG(uint32_t FlagsMask) {
|
||||
|
||||
// Note that the Bfi only considers the bottom bit of the flag, the rest of
|
||||
// the byte is allowed to be garbage.
|
||||
OrderedNode* Flag;
|
||||
Ref Flag;
|
||||
if (FlagOffset == FEXCore::X86State::RFLAG_AF_RAW_LOC) {
|
||||
Flag = LoadAF();
|
||||
} else {
|
||||
@@ -139,10 +139,10 @@ OrderedNode* OpDispatchBuilder::GetPackedRFLAG(uint32_t FlagsMask) {
|
||||
return Original;
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::CalculateOF(uint8_t SrcSize, OrderedNode* Res, OrderedNode* Src1, OrderedNode* Src2, bool Sub) {
|
||||
void OpDispatchBuilder::CalculateOF(uint8_t SrcSize, Ref Res, Ref Src1, Ref Src2, bool Sub) {
|
||||
auto OpSize = SrcSize == 8 ? OpSize::i64Bit : OpSize::i32Bit;
|
||||
uint64_t SignBit = (SrcSize * 8) - 1;
|
||||
OrderedNode* Anded = nullptr;
|
||||
Ref Anded = nullptr;
|
||||
|
||||
// For add, OF is set iff the sources have the same sign but the destination
|
||||
// sign differs. If we know a source sign, we can simplify the expression: if
|
||||
@@ -175,7 +175,7 @@ void OpDispatchBuilder::CalculateOF(uint8_t SrcSize, OrderedNode* Res, OrderedNo
|
||||
SetRFLAG<FEXCore::X86State::RFLAG_OF_RAW_LOC>(Anded, SrcSize * 8 - 1, true);
|
||||
}
|
||||
|
||||
OrderedNode* OpDispatchBuilder::LoadPFRaw(bool Invert) {
|
||||
Ref OpDispatchBuilder::LoadPFRaw(bool Invert) {
|
||||
// Read the stored byte. This is the original result (up to 64-bits), it needs
|
||||
// parity calculated.
|
||||
auto Result = GetRFLAG(FEXCore::X86State::RFLAG_PF_RAW_LOC);
|
||||
@@ -193,7 +193,7 @@ OrderedNode* OpDispatchBuilder::LoadPFRaw(bool Invert) {
|
||||
return Result;
|
||||
}
|
||||
|
||||
OrderedNode* OpDispatchBuilder::LoadAF() {
|
||||
Ref OpDispatchBuilder::LoadAF() {
|
||||
// Read the stored value. This is the XOR of the arguments.
|
||||
auto AFWord = GetRFLAG(FEXCore::X86State::RFLAG_AF_RAW_LOC);
|
||||
|
||||
@@ -214,11 +214,11 @@ void OpDispatchBuilder::FixupAF() {
|
||||
auto PFRaw = GetRFLAG(FEXCore::X86State::RFLAG_PF_RAW_LOC);
|
||||
auto AFRaw = GetRFLAG(FEXCore::X86State::RFLAG_AF_RAW_LOC);
|
||||
|
||||
OrderedNode* XorRes = _Xor(OpSize::i32Bit, AFRaw, PFRaw);
|
||||
Ref XorRes = _Xor(OpSize::i32Bit, AFRaw, PFRaw);
|
||||
SetRFLAG<FEXCore::X86State::RFLAG_AF_RAW_LOC>(XorRes);
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::SetAFAndFixup(OrderedNode* AF) {
|
||||
void OpDispatchBuilder::SetAFAndFixup(Ref AF) {
|
||||
// We have a value of AF, we shift into AF[4]. We need to fixup AF[4] so that
|
||||
// we get the right value when we XOR in PF[4] later. The easiest solution is
|
||||
// to XOR by PF[4], since:
|
||||
@@ -227,16 +227,16 @@ void OpDispatchBuilder::SetAFAndFixup(OrderedNode* AF) {
|
||||
|
||||
auto PFRaw = GetRFLAG(FEXCore::X86State::RFLAG_PF_RAW_LOC);
|
||||
|
||||
OrderedNode* XorRes = _XorShift(OpSize::i32Bit, PFRaw, AF, ShiftType::LSL, 4);
|
||||
Ref XorRes = _XorShift(OpSize::i32Bit, PFRaw, AF, ShiftType::LSL, 4);
|
||||
SetRFLAG<FEXCore::X86State::RFLAG_AF_RAW_LOC>(XorRes);
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::CalculatePF(OrderedNode* Res) {
|
||||
void OpDispatchBuilder::CalculatePF(Ref Res) {
|
||||
// Calculation is entirely deferred until load, just store the 8-bit result.
|
||||
SetRFLAG<FEXCore::X86State::RFLAG_PF_RAW_LOC>(Res);
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::CalculateAF(OrderedNode* Src1, OrderedNode* Src2) {
|
||||
void OpDispatchBuilder::CalculateAF(Ref Src1, Ref Src2) {
|
||||
// We only care about bit 4 in the subsequent XOR. If we'll XOR with 0,
|
||||
// there's no sense XOR'ing at all. If we'll XOR with 1, that's just
|
||||
// inverting.
|
||||
@@ -254,7 +254,7 @@ void OpDispatchBuilder::CalculateAF(OrderedNode* Src1, OrderedNode* Src2) {
|
||||
// We store the XOR of the arguments. At read time, we XOR with the
|
||||
// appropriate bit of the result (available as the PF flag) and extract the
|
||||
// appropriate bit.
|
||||
OrderedNode* XorRes = _Xor(OpSize::i32Bit, Src1, Src2);
|
||||
Ref XorRes = _Xor(OpSize::i32Bit, Src1, Src2);
|
||||
SetRFLAG<FEXCore::X86State::RFLAG_AF_RAW_LOC>(XorRes);
|
||||
}
|
||||
|
||||
@@ -328,11 +328,11 @@ void OpDispatchBuilder::CalculateDeferredFlags(uint32_t FlagsToCalculateMask) {
|
||||
NZCVDirty = false;
|
||||
}
|
||||
|
||||
OrderedNode* OpDispatchBuilder::CalculateFlags_ADC(uint8_t SrcSize, OrderedNode* Src1, OrderedNode* Src2) {
|
||||
Ref OpDispatchBuilder::CalculateFlags_ADC(uint8_t SrcSize, Ref Src1, Ref Src2) {
|
||||
auto Zero = _Constant(0);
|
||||
auto One = _Constant(1);
|
||||
auto OpSize = SrcSize == 8 ? OpSize::i64Bit : OpSize::i32Bit;
|
||||
OrderedNode* Res;
|
||||
Ref Res;
|
||||
|
||||
CalculateAF(Src1, Src2);
|
||||
|
||||
@@ -345,7 +345,7 @@ OrderedNode* OpDispatchBuilder::CalculateFlags_ADC(uint8_t SrcSize, OrderedNode*
|
||||
|
||||
// Note that we do not extend Src2PlusCF, since we depend on proper
|
||||
// 32-bit arithmetic to correctly handle the Src2 = 0xffff case.
|
||||
OrderedNode* Src2PlusCF = _Adc(OpSize, _Constant(0), Src2);
|
||||
Ref Src2PlusCF = _Adc(OpSize, _Constant(0), Src2);
|
||||
|
||||
// Need to zero-extend for the comparison.
|
||||
Res = _Add(OpSize, Src1, Src2PlusCF);
|
||||
@@ -363,14 +363,14 @@ OrderedNode* OpDispatchBuilder::CalculateFlags_ADC(uint8_t SrcSize, OrderedNode*
|
||||
return Res;
|
||||
}
|
||||
|
||||
OrderedNode* OpDispatchBuilder::CalculateFlags_SBB(uint8_t SrcSize, OrderedNode* Src1, OrderedNode* Src2) {
|
||||
Ref OpDispatchBuilder::CalculateFlags_SBB(uint8_t SrcSize, Ref Src1, Ref Src2) {
|
||||
auto Zero = _Constant(0);
|
||||
auto One = _Constant(1);
|
||||
auto OpSize = SrcSize == 8 ? OpSize::i64Bit : OpSize::i32Bit;
|
||||
|
||||
CalculateAF(Src1, Src2);
|
||||
|
||||
OrderedNode* Res;
|
||||
Ref Res;
|
||||
if (SrcSize >= 4) {
|
||||
// Rectify input carry
|
||||
CarryInvert();
|
||||
@@ -402,7 +402,7 @@ OrderedNode* OpDispatchBuilder::CalculateFlags_SBB(uint8_t SrcSize, OrderedNode*
|
||||
return Res;
|
||||
}
|
||||
|
||||
OrderedNode* OpDispatchBuilder::CalculateFlags_SUB(uint8_t SrcSize, OrderedNode* Src1, OrderedNode* Src2, bool UpdateCF) {
|
||||
Ref OpDispatchBuilder::CalculateFlags_SUB(uint8_t SrcSize, Ref Src1, Ref Src2, bool UpdateCF) {
|
||||
// Stash CF before stomping over it
|
||||
auto OldCF = UpdateCF ? nullptr : GetRFLAG(FEXCore::X86State::RFLAG_CF_RAW_LOC);
|
||||
|
||||
@@ -410,7 +410,7 @@ OrderedNode* OpDispatchBuilder::CalculateFlags_SUB(uint8_t SrcSize, OrderedNode*
|
||||
|
||||
CalculateAF(Src1, Src2);
|
||||
|
||||
OrderedNode* Res;
|
||||
Ref Res;
|
||||
if (SrcSize >= 4) {
|
||||
Res = _SubWithFlags(IR::SizeToOpSize(SrcSize), Src1, Src2);
|
||||
} else {
|
||||
@@ -431,7 +431,7 @@ OrderedNode* OpDispatchBuilder::CalculateFlags_SUB(uint8_t SrcSize, OrderedNode*
|
||||
return Res;
|
||||
}
|
||||
|
||||
OrderedNode* OpDispatchBuilder::CalculateFlags_ADD(uint8_t SrcSize, OrderedNode* Src1, OrderedNode* Src2, bool UpdateCF) {
|
||||
Ref OpDispatchBuilder::CalculateFlags_ADD(uint8_t SrcSize, Ref Src1, Ref Src2, bool UpdateCF) {
|
||||
// Stash CF before stomping over it
|
||||
auto OldCF = UpdateCF ? nullptr : GetRFLAG(FEXCore::X86State::RFLAG_CF_RAW_LOC);
|
||||
|
||||
@@ -439,7 +439,7 @@ OrderedNode* OpDispatchBuilder::CalculateFlags_ADD(uint8_t SrcSize, OrderedNode*
|
||||
|
||||
CalculateAF(Src1, Src2);
|
||||
|
||||
OrderedNode* Res;
|
||||
Ref Res;
|
||||
if (SrcSize >= 4) {
|
||||
Res = _AddWithFlags(IR::SizeToOpSize(SrcSize), Src1, Src2);
|
||||
} else {
|
||||
@@ -457,59 +457,39 @@ OrderedNode* OpDispatchBuilder::CalculateFlags_ADD(uint8_t SrcSize, OrderedNode*
|
||||
return Res;
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::CalculateFlags_MUL(uint8_t SrcSize, OrderedNode* Res, OrderedNode* High) {
|
||||
void OpDispatchBuilder::CalculateFlags_MUL(uint8_t SrcSize, Ref Res, Ref High) {
|
||||
HandleNZCVWrite();
|
||||
InvalidatePF_AF();
|
||||
|
||||
// PF/AF/ZF/SF
|
||||
// Undefined
|
||||
{
|
||||
_InvalidateFlags(1 << X86State::RFLAG_PF_RAW_LOC);
|
||||
_InvalidateFlags(1 << X86State::RFLAG_AF_RAW_LOC);
|
||||
}
|
||||
// CF and OF are set if the result of the operation can't be fit in to the destination register
|
||||
// If the value can fit then the top bits will be zero
|
||||
auto SignBit = _Sbfe(OpSize::i64Bit, 1, SrcSize * 8 - 1, Res);
|
||||
_SubNZCV(OpSize::i64Bit, High, SignBit);
|
||||
|
||||
// CF/OF
|
||||
{
|
||||
// CF and OF are set if the result of the operation can't be fit in to the destination register
|
||||
// If the value can fit then the top bits will be zero
|
||||
auto SignBit = _Sbfe(OpSize::i64Bit, 1, SrcSize * 8 - 1, Res);
|
||||
_SubNZCV(OpSize::i64Bit, High, SignBit);
|
||||
|
||||
// If High = SignBit, then sets to nZcv. Else sets to nzCV. Since SF/ZF
|
||||
// undefined, this does what we need.
|
||||
auto Zero = _Constant(0);
|
||||
_CondAddNZCV(OpSize::i64Bit, Zero, Zero, CondClassType {COND_EQ}, 0x3 /* nzCV */);
|
||||
}
|
||||
// If High = SignBit, then sets to nZcv. Else sets to nzCV. Since SF/ZF
|
||||
// undefined, this does what we need.
|
||||
auto Zero = _Constant(0);
|
||||
_CondAddNZCV(OpSize::i64Bit, Zero, Zero, CondClassType {COND_EQ}, 0x3 /* nzCV */);
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::CalculateFlags_UMUL(OrderedNode* High) {
|
||||
void OpDispatchBuilder::CalculateFlags_UMUL(Ref High) {
|
||||
HandleNZCVWrite();
|
||||
InvalidatePF_AF();
|
||||
|
||||
auto Zero = _Constant(0);
|
||||
OpSize Size = IR::SizeToOpSize(GetOpSize(High));
|
||||
|
||||
// AF/SF/PF/ZF
|
||||
// Undefined
|
||||
{
|
||||
_InvalidateFlags(1 << X86State::RFLAG_PF_RAW_LOC);
|
||||
_InvalidateFlags(1 << X86State::RFLAG_AF_RAW_LOC);
|
||||
}
|
||||
// CF and OF are set if the result of the operation can't be fit in to the destination register
|
||||
// The result register will be all zero if it can't fit due to how multiplication behaves
|
||||
_SubNZCV(Size, High, Zero);
|
||||
|
||||
// CF/OF
|
||||
{
|
||||
// CF and OF are set if the result of the operation can't be fit in to the destination register
|
||||
// The result register will be all zero if it can't fit due to how multiplication behaves
|
||||
_SubNZCV(Size, High, Zero);
|
||||
|
||||
// If High = 0, then sets to nZcv. Else sets to nzCV. Since SF/ZF undefined,
|
||||
// this does what we need.
|
||||
_CondAddNZCV(Size, Zero, Zero, CondClassType {COND_EQ}, 0x3 /* nzCV */);
|
||||
}
|
||||
// If High = 0, then sets to nZcv. Else sets to nzCV. Since SF/ZF undefined,
|
||||
// this does what we need.
|
||||
_CondAddNZCV(Size, Zero, Zero, CondClassType {COND_EQ}, 0x3 /* nzCV */);
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::CalculateFlags_Logical(uint8_t SrcSize, OrderedNode* Res, OrderedNode* Src1, OrderedNode* Src2) {
|
||||
// AF
|
||||
// Undefined
|
||||
_InvalidateFlags(1 << X86State::RFLAG_AF_RAW_LOC);
|
||||
void OpDispatchBuilder::CalculateFlags_Logical(uint8_t SrcSize, Ref Res, Ref Src1, Ref Src2) {
|
||||
InvalidateAF();
|
||||
|
||||
CalculatePF(Res);
|
||||
|
||||
@@ -517,7 +497,7 @@ void OpDispatchBuilder::CalculateFlags_Logical(uint8_t SrcSize, OrderedNode* Res
|
||||
SetNZ_ZeroCV(SrcSize, Res);
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::CalculateFlags_ShiftLeftImmediate(uint8_t SrcSize, OrderedNode* UnmaskedRes, OrderedNode* Src1, uint64_t Shift) {
|
||||
void OpDispatchBuilder::CalculateFlags_ShiftLeftImmediate(uint8_t SrcSize, Ref UnmaskedRes, Ref Src1, uint64_t Shift) {
|
||||
// No flags changed if shift is zero
|
||||
if (Shift == 0) {
|
||||
return;
|
||||
@@ -538,10 +518,7 @@ void OpDispatchBuilder::CalculateFlags_ShiftLeftImmediate(uint8_t SrcSize, Order
|
||||
}
|
||||
|
||||
CalculatePF(UnmaskedRes);
|
||||
|
||||
// AF
|
||||
// Undefined
|
||||
_InvalidateFlags(1 << X86State::RFLAG_AF_RAW_LOC);
|
||||
InvalidateAF();
|
||||
|
||||
// OF
|
||||
// In the case of left shift. OF is only set from the result of <Top Source Bit> XOR <Top Result Bit>
|
||||
@@ -553,7 +530,7 @@ void OpDispatchBuilder::CalculateFlags_ShiftLeftImmediate(uint8_t SrcSize, Order
|
||||
}
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::CalculateFlags_SignShiftRightImmediate(uint8_t SrcSize, OrderedNode* Res, OrderedNode* Src1, uint64_t Shift) {
|
||||
void OpDispatchBuilder::CalculateFlags_SignShiftRightImmediate(uint8_t SrcSize, Ref Res, Ref Src1, uint64_t Shift) {
|
||||
// No flags changed if shift is zero
|
||||
if (Shift == 0) {
|
||||
return;
|
||||
@@ -568,10 +545,7 @@ void OpDispatchBuilder::CalculateFlags_SignShiftRightImmediate(uint8_t SrcSize,
|
||||
}
|
||||
|
||||
CalculatePF(Res);
|
||||
|
||||
// AF
|
||||
// Undefined
|
||||
_InvalidateFlags(1 << X86State::RFLAG_AF_RAW_LOC);
|
||||
InvalidateAF();
|
||||
|
||||
// OF
|
||||
// Only defined when Shift is 1 else undefined. Only is set if the top bit was set to 1 when
|
||||
@@ -579,7 +553,7 @@ void OpDispatchBuilder::CalculateFlags_SignShiftRightImmediate(uint8_t SrcSize,
|
||||
// already zeroed there's nothing to do here.
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::CalculateFlags_ShiftRightImmediateCommon(uint8_t SrcSize, OrderedNode* Res, OrderedNode* Src1, uint64_t Shift) {
|
||||
void OpDispatchBuilder::CalculateFlags_ShiftRightImmediateCommon(uint8_t SrcSize, Ref Res, Ref Src1, uint64_t Shift) {
|
||||
// Set SF and PF. Clobbers OF, but OF only defined for Shift = 1 where it is
|
||||
// set below.
|
||||
SetNZ_ZeroCV(SrcSize, Res);
|
||||
@@ -591,13 +565,10 @@ void OpDispatchBuilder::CalculateFlags_ShiftRightImmediateCommon(uint8_t SrcSize
|
||||
}
|
||||
|
||||
CalculatePF(Res);
|
||||
|
||||
// AF
|
||||
// Undefined
|
||||
_InvalidateFlags(1 << X86State::RFLAG_AF_RAW_LOC);
|
||||
InvalidateAF();
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::CalculateFlags_ShiftRightImmediate(uint8_t SrcSize, OrderedNode* Res, OrderedNode* Src1, uint64_t Shift) {
|
||||
void OpDispatchBuilder::CalculateFlags_ShiftRightImmediate(uint8_t SrcSize, Ref Res, Ref Src1, uint64_t Shift) {
|
||||
// No flags changed if shift is zero
|
||||
if (Shift == 0) {
|
||||
return;
|
||||
@@ -615,7 +586,7 @@ void OpDispatchBuilder::CalculateFlags_ShiftRightImmediate(uint8_t SrcSize, Orde
|
||||
}
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::CalculateFlags_ShiftRightDoubleImmediate(uint8_t SrcSize, OrderedNode* Res, OrderedNode* Src1, uint64_t Shift) {
|
||||
void OpDispatchBuilder::CalculateFlags_ShiftRightDoubleImmediate(uint8_t SrcSize, Ref Res, Ref Src1, uint64_t Shift) {
|
||||
// No flags changed if shift is zero
|
||||
if (Shift == 0) {
|
||||
return;
|
||||
@@ -636,32 +607,28 @@ void OpDispatchBuilder::CalculateFlags_ShiftRightDoubleImmediate(uint8_t SrcSize
|
||||
}
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::CalculateFlags_BEXTR(OrderedNode* Src) {
|
||||
void OpDispatchBuilder::CalculateFlags_BEXTR(Ref Src) {
|
||||
// ZF is set properly. CF and OF are defined as being set to zero. SF, PF, and
|
||||
// AF are undefined.
|
||||
SetNZ_ZeroCV(GetOpSize(Src), Src);
|
||||
|
||||
_InvalidateFlags((1UL << X86State::RFLAG_PF_RAW_LOC) | (1UL << X86State::RFLAG_AF_RAW_LOC));
|
||||
InvalidatePF_AF();
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::CalculateFlags_BLSI(uint8_t SrcSize, OrderedNode* Result) {
|
||||
void OpDispatchBuilder::CalculateFlags_BLSI(uint8_t SrcSize, Ref Result) {
|
||||
// CF is cleared if Src is zero, otherwise it's set. However, Src is zero iff
|
||||
// Result is zero, so we can test the result instead. So, CF is just the
|
||||
// inverted ZF.
|
||||
//
|
||||
// ZF/SF/OF set as usual.
|
||||
SetNZ_ZeroCV(SrcSize, Result);
|
||||
InvalidatePF_AF();
|
||||
|
||||
auto CFOp = GetRFLAG(X86State::RFLAG_ZF_RAW_LOC, true /* Invert */);
|
||||
SetRFLAG<X86State::RFLAG_CF_RAW_LOC>(CFOp);
|
||||
|
||||
// PF/AF undefined
|
||||
_InvalidateFlags((1UL << X86State::RFLAG_PF_RAW_LOC) | (1UL << X86State::RFLAG_AF_RAW_LOC));
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::CalculateFlags_BLSMSK(uint8_t SrcSize, OrderedNode* Result, OrderedNode* Src) {
|
||||
// PF/AF undefined
|
||||
_InvalidateFlags((1UL << X86State::RFLAG_PF_RAW_LOC) | (1UL << X86State::RFLAG_AF_RAW_LOC));
|
||||
void OpDispatchBuilder::CalculateFlags_BLSMSK(uint8_t SrcSize, Ref Result, Ref Src) {
|
||||
InvalidatePF_AF();
|
||||
|
||||
// CF set according to the Src
|
||||
auto Zero = _Constant(0);
|
||||
@@ -674,19 +641,17 @@ void OpDispatchBuilder::CalculateFlags_BLSMSK(uint8_t SrcSize, OrderedNode* Resu
|
||||
SetRFLAG<X86State::RFLAG_CF_RAW_LOC>(CFOp);
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::CalculateFlags_BLSR(uint8_t SrcSize, OrderedNode* Result, OrderedNode* Src) {
|
||||
void OpDispatchBuilder::CalculateFlags_BLSR(uint8_t SrcSize, Ref Result, Ref Src) {
|
||||
auto Zero = _Constant(0);
|
||||
auto One = _Constant(1);
|
||||
auto CFOp = _Select(IR::COND_EQ, Src, Zero, One, Zero);
|
||||
|
||||
SetNZ_ZeroCV(SrcSize, Result);
|
||||
SetRFLAG<X86State::RFLAG_CF_RAW_LOC>(CFOp);
|
||||
|
||||
// PF/AF undefined
|
||||
_InvalidateFlags((1UL << X86State::RFLAG_PF_RAW_LOC) | (1UL << X86State::RFLAG_AF_RAW_LOC));
|
||||
InvalidatePF_AF();
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::CalculateFlags_POPCOUNT(OrderedNode* Result) {
|
||||
void OpDispatchBuilder::CalculateFlags_POPCOUNT(Ref Result) {
|
||||
// We need to set ZF while clearing the rest of NZCV. The result of a popcount
|
||||
// is in the range [0, 63]. In particular, it is always positive. So a
|
||||
// combined NZ test will correctly zero SF/CF/OF while setting ZF.
|
||||
@@ -694,15 +659,13 @@ void OpDispatchBuilder::CalculateFlags_POPCOUNT(OrderedNode* Result) {
|
||||
ZeroPF_AF();
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::CalculateFlags_BZHI(uint8_t SrcSize, OrderedNode* Result, OrderedNode* Src) {
|
||||
// PF/AF undefined
|
||||
_InvalidateFlags((1UL << X86State::RFLAG_PF_RAW_LOC) | (1UL << X86State::RFLAG_AF_RAW_LOC));
|
||||
|
||||
void OpDispatchBuilder::CalculateFlags_BZHI(uint8_t SrcSize, Ref Result, Ref Src) {
|
||||
InvalidatePF_AF();
|
||||
SetNZ_ZeroCV(SrcSize, Result);
|
||||
SetRFLAG<X86State::RFLAG_CF_RAW_LOC>(Src);
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::CalculateFlags_ZCNT(uint8_t SrcSize, OrderedNode* Result) {
|
||||
void OpDispatchBuilder::CalculateFlags_ZCNT(uint8_t SrcSize, Ref Result) {
|
||||
// OF, SF, AF, PF all undefined
|
||||
// Test ZF of result, SF is undefined so this is ok.
|
||||
SetNZ_ZeroCV(SrcSize, Result);
|
||||
@@ -714,7 +677,7 @@ void OpDispatchBuilder::CalculateFlags_ZCNT(uint8_t SrcSize, OrderedNode* Result
|
||||
SetRFLAG<FEXCore::X86State::RFLAG_CF_RAW_LOC>(Result, CarryBit);
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::CalculateFlags_RDRAND(OrderedNode* Src) {
|
||||
void OpDispatchBuilder::CalculateFlags_RDRAND(Ref Src) {
|
||||
// OF, SF, ZF, AF, PF all zero
|
||||
ZeroNZCV();
|
||||
ZeroPF_AF();
|
||||
|
||||
File diff suppressed because it is too large.
Load diff
@@ -23,45 +23,45 @@ class OrderedNode;
|
||||
|
||||
#define OpcodeArgs [[maybe_unused]] FEXCore::X86Tables::DecodedOp Op
|
||||
|
||||
OrderedNode* OpDispatchBuilder::GetX87Top() {
|
||||
Ref OpDispatchBuilder::GetX87Top() {
|
||||
// Yes, we are storing 3 bits in a single flag register.
|
||||
// Deal with it
|
||||
return _LoadContext(1, GPRClass, offsetof(FEXCore::Core::CPUState, flags) + FEXCore::X86State::X87FLAG_TOP_LOC);
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::SetX87ValidTag(OrderedNode* Value, bool Valid) {
|
||||
void OpDispatchBuilder::SetX87ValidTag(Ref Value, bool Valid) {
|
||||
// if we are popping then we must first mark this location as empty
|
||||
OrderedNode* AbridgedFTW = _LoadContext(1, GPRClass, offsetof(FEXCore::Core::CPUState, AbridgedFTW));
|
||||
OrderedNode* RegMask = _Lshl(OpSize::i32Bit, _Constant(1), Value);
|
||||
OrderedNode* NewAbridgedFTW = Valid ? _Or(OpSize::i32Bit, AbridgedFTW, RegMask) : _Andn(OpSize::i32Bit, AbridgedFTW, RegMask);
|
||||
Ref AbridgedFTW = _LoadContext(1, GPRClass, offsetof(FEXCore::Core::CPUState, AbridgedFTW));
|
||||
Ref RegMask = _Lshl(OpSize::i32Bit, _Constant(1), Value);
|
||||
Ref NewAbridgedFTW = Valid ? _Or(OpSize::i32Bit, AbridgedFTW, RegMask) : _Andn(OpSize::i32Bit, AbridgedFTW, RegMask);
|
||||
_StoreContext(1, GPRClass, NewAbridgedFTW, offsetof(FEXCore::Core::CPUState, AbridgedFTW));
|
||||
}
|
||||
|
||||
OrderedNode* OpDispatchBuilder::GetX87ValidTag(OrderedNode* Value) {
|
||||
OrderedNode* AbridgedFTW = _LoadContext(1, GPRClass, offsetof(FEXCore::Core::CPUState, AbridgedFTW));
|
||||
Ref OpDispatchBuilder::GetX87ValidTag(Ref Value) {
|
||||
Ref AbridgedFTW = _LoadContext(1, GPRClass, offsetof(FEXCore::Core::CPUState, AbridgedFTW));
|
||||
return _And(OpSize::i32Bit, _Lshr(OpSize::i32Bit, AbridgedFTW, Value), _Constant(1));
|
||||
}
|
||||
|
||||
OrderedNode* OpDispatchBuilder::GetX87Tag(OrderedNode* Value, OrderedNode* AbridgedFTW) {
|
||||
OrderedNode* RegValid = _And(OpSize::i32Bit, _Lshr(OpSize::i32Bit, AbridgedFTW, Value), _Constant(1));
|
||||
OrderedNode* X87Empty = _Constant(static_cast<uint8_t>(FPState::X87Tag::Empty));
|
||||
OrderedNode* X87Valid = _Constant(static_cast<uint8_t>(FPState::X87Tag::Valid));
|
||||
Ref OpDispatchBuilder::GetX87Tag(Ref Value, Ref AbridgedFTW) {
|
||||
Ref RegValid = _And(OpSize::i32Bit, _Lshr(OpSize::i32Bit, AbridgedFTW, Value), _Constant(1));
|
||||
Ref X87Empty = _Constant(static_cast<uint8_t>(FPState::X87Tag::Empty));
|
||||
Ref X87Valid = _Constant(static_cast<uint8_t>(FPState::X87Tag::Valid));
|
||||
|
||||
return _Select(FEXCore::IR::COND_EQ, RegValid, _Constant(0), X87Empty, X87Valid);
|
||||
}
|
||||
|
||||
OrderedNode* OpDispatchBuilder::GetX87Tag(OrderedNode* Value) {
|
||||
OrderedNode* AbridgedFTW = _LoadContext(1, GPRClass, offsetof(FEXCore::Core::CPUState, AbridgedFTW));
|
||||
Ref OpDispatchBuilder::GetX87Tag(Ref Value) {
|
||||
Ref AbridgedFTW = _LoadContext(1, GPRClass, offsetof(FEXCore::Core::CPUState, AbridgedFTW));
|
||||
return GetX87Tag(Value, AbridgedFTW);
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::SetX87FTW(OrderedNode* FTW) {
|
||||
OrderedNode* X87Empty = _Constant(static_cast<uint8_t>(FPState::X87Tag::Empty));
|
||||
OrderedNode* NewAbridgedFTW;
|
||||
void OpDispatchBuilder::SetX87FTW(Ref FTW) {
|
||||
Ref X87Empty = _Constant(static_cast<uint8_t>(FPState::X87Tag::Empty));
|
||||
Ref NewAbridgedFTW;
|
||||
|
||||
for (int i = 0; i < 8; i++) {
|
||||
OrderedNode* RegTag = _Bfe(OpSize::i32Bit, 2, i * 2, FTW);
|
||||
OrderedNode* RegValid = _Select(FEXCore::IR::COND_NEQ, RegTag, X87Empty, _Constant(1), _Constant(0));
|
||||
Ref RegTag = _Bfe(OpSize::i32Bit, 2, i * 2, FTW);
|
||||
Ref RegValid = _Select(FEXCore::IR::COND_NEQ, RegTag, X87Empty, _Constant(1), _Constant(0));
|
||||
|
||||
if (i) {
|
||||
NewAbridgedFTW = _Orlshl(OpSize::i32Bit, NewAbridgedFTW, RegValid, i);
|
||||
@@ -73,9 +73,9 @@ void OpDispatchBuilder::SetX87FTW(OrderedNode* FTW) {
|
||||
_StoreContext(1, GPRClass, NewAbridgedFTW, offsetof(FEXCore::Core::CPUState, AbridgedFTW));
|
||||
}
|
||||
|
||||
OrderedNode* OpDispatchBuilder::GetX87FTW() {
|
||||
OrderedNode* AbridgedFTW = _LoadContext(1, GPRClass, offsetof(FEXCore::Core::CPUState, AbridgedFTW));
|
||||
OrderedNode* FTW = _Constant(0);
|
||||
Ref OpDispatchBuilder::GetX87FTW() {
|
||||
Ref AbridgedFTW = _LoadContext(1, GPRClass, offsetof(FEXCore::Core::CPUState, AbridgedFTW));
|
||||
Ref FTW = _Constant(0);
|
||||
|
||||
for (int i = 0; i < 8; i++) {
|
||||
const auto RegTag = GetX87Tag(_Constant(i), AbridgedFTW);
|
||||
@@ -85,13 +85,13 @@ OrderedNode* OpDispatchBuilder::GetX87FTW() {
|
||||
return FTW;
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::SetX87Top(OrderedNode* Value) {
|
||||
void OpDispatchBuilder::SetX87Top(Ref Value) {
|
||||
_StoreContext(1, GPRClass, Value, offsetof(FEXCore::Core::CPUState, flags) + FEXCore::X86State::X87FLAG_TOP_LOC);
|
||||
}
|
||||
|
||||
OrderedNode* OpDispatchBuilder::ReconstructFSW() {
|
||||
Ref OpDispatchBuilder::ReconstructFSW() {
|
||||
// We must construct the FSW from our various bits
|
||||
OrderedNode* FSW = _Constant(0);
|
||||
Ref FSW = _Constant(0);
|
||||
auto Top = GetX87Top();
|
||||
FSW = _Bfi(OpSize::i64Bit, 3, 11, FSW, Top);
|
||||
|
||||
@@ -107,7 +107,7 @@ OrderedNode* OpDispatchBuilder::ReconstructFSW() {
|
||||
return FSW;
|
||||
}
|
||||
|
||||
OrderedNode* OpDispatchBuilder::ReconstructX87StateFromFSW(OrderedNode* FSW) {
|
||||
Ref OpDispatchBuilder::ReconstructX87StateFromFSW(Ref FSW) {
|
||||
auto Top = _Bfe(OpSize::i32Bit, 3, 11, FSW);
|
||||
SetX87Top(Top);
|
||||
|
||||
@@ -131,7 +131,7 @@ void OpDispatchBuilder::FLD(OpcodeArgs) {
|
||||
|
||||
size_t read_width = (width == 80) ? 16 : width / 8;
|
||||
|
||||
OrderedNode* data {};
|
||||
Ref data {};
|
||||
|
||||
if (!Op->Src[0].IsNone()) {
|
||||
// Read from memory
|
||||
@@ -142,7 +142,7 @@ void OpDispatchBuilder::FLD(OpcodeArgs) {
|
||||
data = _And(OpSize::i32Bit, _Add(OpSize::i32Bit, orig_top, offset), mask);
|
||||
data = _LoadContextIndexed(data, 16, MMBaseOffset(), 16, FPRClass);
|
||||
}
|
||||
OrderedNode* converted = data;
|
||||
Ref converted = data;
|
||||
|
||||
// Convert to 80bit float
|
||||
if constexpr (width == 32 || width == 64) {
|
||||
@@ -170,8 +170,8 @@ void OpDispatchBuilder::FBLD(OpcodeArgs) {
|
||||
SetX87Top(top);
|
||||
|
||||
// Read from memory
|
||||
OrderedNode* data = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], 16, Op->Flags);
|
||||
OrderedNode* converted = _F80BCDLoad(data);
|
||||
Ref data = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], 16, Op->Flags);
|
||||
Ref converted = _F80BCDLoad(data);
|
||||
_StoreContextIndexed(converted, top, 16, MMBaseOffset(), 16, FPRClass);
|
||||
}
|
||||
|
||||
@@ -179,7 +179,7 @@ void OpDispatchBuilder::FBSTP(OpcodeArgs) {
|
||||
auto orig_top = GetX87Top();
|
||||
auto data = _LoadContextIndexed(orig_top, 16, MMBaseOffset(), 16, FPRClass);
|
||||
|
||||
OrderedNode* converted = _F80BCDStore(data);
|
||||
Ref converted = _F80BCDStore(data);
|
||||
|
||||
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, converted, 10, 1);
|
||||
|
||||
@@ -197,7 +197,7 @@ void OpDispatchBuilder::FLD_Const(OpcodeArgs) {
|
||||
SetX87ValidTag(top, true);
|
||||
SetX87Top(top);
|
||||
|
||||
OrderedNode* data = LoadAndCacheNamedVectorConstant(16, constant);
|
||||
Ref data = LoadAndCacheNamedVectorConstant(16, constant);
|
||||
|
||||
// Write to ST[TOP]
|
||||
_StoreContextIndexed(data, top, 16, MMBaseOffset(), 16, FPRClass);
|
||||
@@ -247,7 +247,7 @@ void OpDispatchBuilder::FILD(OpcodeArgs) {
|
||||
auto upper = _Or(OpSize::i64Bit, sign, zeroed_exponent);
|
||||
|
||||
|
||||
OrderedNode* converted = _VCastFromGPR(16, 8, shifted);
|
||||
Ref converted = _VCastFromGPR(16, 8, shifted);
|
||||
converted = _VInsElement(16, 8, 1, 0, converted, _VCastFromGPR(16, 8, upper));
|
||||
|
||||
// Write to ST[TOP]
|
||||
@@ -283,7 +283,7 @@ void OpDispatchBuilder::FIST(OpcodeArgs) {
|
||||
auto Size = GetSrcSize(Op);
|
||||
|
||||
auto orig_top = GetX87Top();
|
||||
OrderedNode* data = _LoadContextIndexed(orig_top, 16, MMBaseOffset(), 16, FPRClass);
|
||||
Ref data = _LoadContextIndexed(orig_top, 16, MMBaseOffset(), 16, FPRClass);
|
||||
data = _F80CVTInt(Size, data, Truncate);
|
||||
|
||||
StoreResult_WithOpSize(GPRClass, Op, Op->Dest, data, Size, 1);
|
||||
@@ -303,10 +303,10 @@ template void OpDispatchBuilder::FIST<true>(OpcodeArgs);
|
||||
template<size_t width, bool Integer, OpDispatchBuilder::OpResult ResInST0>
|
||||
void OpDispatchBuilder::FADD(OpcodeArgs) {
|
||||
auto top = GetX87Top();
|
||||
OrderedNode* StackLocation = top;
|
||||
Ref StackLocation = top;
|
||||
|
||||
OrderedNode* arg {};
|
||||
OrderedNode* b {};
|
||||
Ref arg {};
|
||||
Ref b {};
|
||||
|
||||
auto mask = _Constant(7);
|
||||
|
||||
@@ -357,9 +357,9 @@ template void OpDispatchBuilder::FADD<32, true, OpDispatchBuilder::OpResult::RES
|
||||
template<size_t width, bool Integer, OpDispatchBuilder::OpResult ResInST0>
|
||||
void OpDispatchBuilder::FMUL(OpcodeArgs) {
|
||||
auto top = GetX87Top();
|
||||
OrderedNode* StackLocation = top;
|
||||
OrderedNode* arg {};
|
||||
OrderedNode* b {};
|
||||
Ref StackLocation = top;
|
||||
Ref arg {};
|
||||
Ref b {};
|
||||
|
||||
auto mask = _Constant(7);
|
||||
|
||||
@@ -413,9 +413,9 @@ template void OpDispatchBuilder::FMUL<32, true, OpDispatchBuilder::OpResult::RES
|
||||
template<size_t width, bool Integer, bool reverse, OpDispatchBuilder::OpResult ResInST0>
|
||||
void OpDispatchBuilder::FDIV(OpcodeArgs) {
|
||||
auto top = GetX87Top();
|
||||
OrderedNode* StackLocation = top;
|
||||
OrderedNode* arg {};
|
||||
OrderedNode* b {};
|
||||
Ref StackLocation = top;
|
||||
Ref arg {};
|
||||
Ref b {};
|
||||
|
||||
auto mask = _Constant(7);
|
||||
|
||||
@@ -444,7 +444,7 @@ void OpDispatchBuilder::FDIV(OpcodeArgs) {
|
||||
|
||||
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
|
||||
|
||||
OrderedNode* result {};
|
||||
Ref result {};
|
||||
if constexpr (reverse) {
|
||||
result = _F80Div(b, a);
|
||||
} else {
|
||||
@@ -484,9 +484,9 @@ template void OpDispatchBuilder::FDIV<32, true, true, OpDispatchBuilder::OpResul
|
||||
template<size_t width, bool Integer, bool reverse, OpDispatchBuilder::OpResult ResInST0>
|
||||
void OpDispatchBuilder::FSUB(OpcodeArgs) {
|
||||
auto top = GetX87Top();
|
||||
OrderedNode* StackLocation = top;
|
||||
OrderedNode* arg {};
|
||||
OrderedNode* b {};
|
||||
Ref StackLocation = top;
|
||||
Ref arg {};
|
||||
Ref b {};
|
||||
|
||||
auto mask = _Constant(7);
|
||||
|
||||
@@ -514,7 +514,7 @@ void OpDispatchBuilder::FSUB(OpcodeArgs) {
|
||||
|
||||
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
|
||||
|
||||
OrderedNode* result {};
|
||||
Ref result {};
|
||||
if constexpr (reverse) {
|
||||
result = _F80Sub(b, a);
|
||||
} else {
|
||||
@@ -558,7 +558,7 @@ void OpDispatchBuilder::FCHS(OpcodeArgs) {
|
||||
|
||||
auto low = _Constant(0);
|
||||
auto high = _Constant(0b1'000'0000'0000'0000ULL);
|
||||
OrderedNode* data = _VCastFromGPR(16, 8, low);
|
||||
Ref data = _VCastFromGPR(16, 8, low);
|
||||
data = _VInsGPR(16, 8, 1, data, high);
|
||||
|
||||
auto result = _VXor(16, 1, a, data);
|
||||
@@ -573,7 +573,7 @@ void OpDispatchBuilder::FABS(OpcodeArgs) {
|
||||
|
||||
auto low = _Constant(~0ULL);
|
||||
auto high = _Constant(0b0'111'1111'1111'1111ULL);
|
||||
OrderedNode* data = _VCastFromGPR(16, 8, low);
|
||||
Ref data = _VCastFromGPR(16, 8, low);
|
||||
data = _VInsGPR(16, 8, 1, data, high);
|
||||
|
||||
auto result = _VAnd(16, 1, a, data);
|
||||
@@ -587,13 +587,13 @@ void OpDispatchBuilder::FTST(OpcodeArgs) {
|
||||
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
|
||||
|
||||
auto low = _Constant(0);
|
||||
OrderedNode* data = _VCastFromGPR(16, 8, low);
|
||||
Ref data = _VCastFromGPR(16, 8, low);
|
||||
|
||||
OrderedNode* Res = _F80Cmp(a, data, (1 << FCMP_FLAG_EQ) | (1 << FCMP_FLAG_LT) | (1 << FCMP_FLAG_UNORDERED));
|
||||
Ref Res = _F80Cmp(a, data, (1 << FCMP_FLAG_EQ) | (1 << FCMP_FLAG_LT) | (1 << FCMP_FLAG_UNORDERED));
|
||||
|
||||
OrderedNode* HostFlag_CF = _GetHostFlag(Res, FCMP_FLAG_LT);
|
||||
OrderedNode* HostFlag_ZF = _GetHostFlag(Res, FCMP_FLAG_EQ);
|
||||
OrderedNode* HostFlag_Unordered = _GetHostFlag(Res, FCMP_FLAG_UNORDERED);
|
||||
Ref HostFlag_CF = _GetHostFlag(Res, FCMP_FLAG_LT);
|
||||
Ref HostFlag_ZF = _GetHostFlag(Res, FCMP_FLAG_EQ);
|
||||
Ref HostFlag_Unordered = _GetHostFlag(Res, FCMP_FLAG_UNORDERED);
|
||||
HostFlag_CF = _Or(OpSize::i32Bit, HostFlag_CF, HostFlag_Unordered);
|
||||
HostFlag_ZF = _Or(OpSize::i32Bit, HostFlag_ZF, HostFlag_Unordered);
|
||||
|
||||
@@ -652,8 +652,8 @@ void OpDispatchBuilder::FCOMI(OpcodeArgs) {
|
||||
auto top = GetX87Top();
|
||||
auto mask = _Constant(7);
|
||||
|
||||
OrderedNode* arg {};
|
||||
OrderedNode* b {};
|
||||
Ref arg {};
|
||||
Ref b {};
|
||||
|
||||
if (!Op->Src[0].IsNone()) {
|
||||
// Memory arg
|
||||
@@ -675,11 +675,11 @@ void OpDispatchBuilder::FCOMI(OpcodeArgs) {
|
||||
|
||||
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
|
||||
|
||||
OrderedNode* Res = _F80Cmp(a, b, (1 << FCMP_FLAG_EQ) | (1 << FCMP_FLAG_LT) | (1 << FCMP_FLAG_UNORDERED));
|
||||
Ref Res = _F80Cmp(a, b, (1 << FCMP_FLAG_EQ) | (1 << FCMP_FLAG_LT) | (1 << FCMP_FLAG_UNORDERED));
|
||||
|
||||
OrderedNode* HostFlag_CF = _GetHostFlag(Res, FCMP_FLAG_LT);
|
||||
OrderedNode* HostFlag_ZF = _GetHostFlag(Res, FCMP_FLAG_EQ);
|
||||
OrderedNode* HostFlag_Unordered = _GetHostFlag(Res, FCMP_FLAG_UNORDERED);
|
||||
Ref HostFlag_CF = _GetHostFlag(Res, FCMP_FLAG_LT);
|
||||
Ref HostFlag_ZF = _GetHostFlag(Res, FCMP_FLAG_EQ);
|
||||
Ref HostFlag_Unordered = _GetHostFlag(Res, FCMP_FLAG_UNORDERED);
|
||||
HostFlag_CF = _Or(OpSize::i32Bit, HostFlag_CF, HostFlag_Unordered);
|
||||
HostFlag_ZF = _Or(OpSize::i32Bit, HostFlag_ZF, HostFlag_Unordered);
|
||||
|
||||
@@ -734,7 +734,7 @@ template void OpDispatchBuilder::FCOMI<32, true, OpDispatchBuilder::FCOMIFlags::
|
||||
|
||||
void OpDispatchBuilder::FXCH(OpcodeArgs) {
|
||||
auto top = GetX87Top();
|
||||
OrderedNode* arg;
|
||||
Ref arg;
|
||||
|
||||
auto mask = _Constant(7);
|
||||
|
||||
@@ -745,6 +745,9 @@ void OpDispatchBuilder::FXCH(OpcodeArgs) {
|
||||
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
|
||||
auto b = _LoadContextIndexed(arg, 16, MMBaseOffset(), 16, FPRClass);
|
||||
|
||||
// Set C1 to Zero
|
||||
SetRFLAG<FEXCore::X86State::X87FLAG_C1_LOC>(_Constant(0));
|
||||
|
||||
// Write to ST[TOP]
|
||||
_StoreContextIndexed(b, top, 16, MMBaseOffset(), 16, FPRClass);
|
||||
_StoreContextIndexed(a, arg, 16, MMBaseOffset(), 16, FPRClass);
|
||||
@@ -752,19 +755,20 @@ void OpDispatchBuilder::FXCH(OpcodeArgs) {
|
||||
|
||||
void OpDispatchBuilder::FST(OpcodeArgs) {
|
||||
auto top = GetX87Top();
|
||||
OrderedNode* arg;
|
||||
|
||||
auto mask = _Constant(7);
|
||||
|
||||
// Implicit arg
|
||||
auto offset = _Constant(Op->OP & 7);
|
||||
arg = _And(OpSize::i32Bit, _Add(OpSize::i32Bit, top, offset), mask);
|
||||
Ref arg = _And(OpSize::i32Bit, _Add(OpSize::i32Bit, top, offset), mask);
|
||||
|
||||
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
|
||||
|
||||
// Write to ST[TOP]
|
||||
// Write to ST[i]
|
||||
_StoreContextIndexed(a, arg, 16, MMBaseOffset(), 16, FPRClass);
|
||||
|
||||
// Set Tag for ST[i]
|
||||
SetX87ValidTag(arg, true);
|
||||
|
||||
if ((Op->TableInfo->Flags & X86Tables::InstFlags::FLAGS_POP) != 0) {
|
||||
// if we are popping then we must first mark this location as empty
|
||||
SetX87ValidTag(top, false);
|
||||
@@ -799,7 +803,7 @@ void OpDispatchBuilder::X87BinaryOp(OpcodeArgs) {
|
||||
auto top = GetX87Top();
|
||||
|
||||
auto mask = _Constant(7);
|
||||
OrderedNode* st1 = _And(OpSize::i32Bit, _Add(OpSize::i32Bit, top, _Constant(1)), mask);
|
||||
Ref st1 = _And(OpSize::i32Bit, _Add(OpSize::i32Bit, top, _Constant(1)), mask);
|
||||
|
||||
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
|
||||
st1 = _LoadContextIndexed(st1, 16, MMBaseOffset(), 16, FPRClass);
|
||||
@@ -862,11 +866,11 @@ void OpDispatchBuilder::X87FYL2X(OpcodeArgs) {
|
||||
auto top = _And(OpSize::i32Bit, _Add(OpSize::i32Bit, orig_top, _Constant(1)), _Constant(7));
|
||||
SetX87Top(top);
|
||||
|
||||
OrderedNode* st0 = _LoadContextIndexed(orig_top, 16, MMBaseOffset(), 16, FPRClass);
|
||||
OrderedNode* st1 = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
|
||||
Ref st0 = _LoadContextIndexed(orig_top, 16, MMBaseOffset(), 16, FPRClass);
|
||||
Ref st1 = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
|
||||
|
||||
if (Plus1) {
|
||||
OrderedNode* data = LoadAndCacheNamedVectorConstant(16, NamedVectorConstant::NAMED_VECTOR_X87_ONE);
|
||||
Ref data = LoadAndCacheNamedVectorConstant(16, NamedVectorConstant::NAMED_VECTOR_X87_ONE);
|
||||
st0 = _F80Add(st0, data);
|
||||
}
|
||||
|
||||
@@ -886,7 +890,7 @@ void OpDispatchBuilder::X87TAN(OpcodeArgs) {
|
||||
|
||||
auto result = _F80TAN(a);
|
||||
|
||||
OrderedNode* data = LoadAndCacheNamedVectorConstant(16, NamedVectorConstant::NAMED_VECTOR_X87_ONE);
|
||||
Ref data = LoadAndCacheNamedVectorConstant(16, NamedVectorConstant::NAMED_VECTOR_X87_ONE);
|
||||
|
||||
// TODO: ACCURACY: should check source is in range –2^63 to +2^63
|
||||
SetRFLAG<FEXCore::X86State::X87FLAG_C2_LOC>(_Constant(0));
|
||||
@@ -904,7 +908,7 @@ void OpDispatchBuilder::X87ATAN(OpcodeArgs) {
|
||||
SetX87Top(top);
|
||||
|
||||
auto a = _LoadContextIndexed(orig_top, 16, MMBaseOffset(), 16, FPRClass);
|
||||
OrderedNode* st1 = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
|
||||
Ref st1 = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
|
||||
|
||||
auto result = _F80ATAN(st1, a);
|
||||
|
||||
@@ -914,19 +918,17 @@ void OpDispatchBuilder::X87ATAN(OpcodeArgs) {
|
||||
|
||||
void OpDispatchBuilder::X87LDENV(OpcodeArgs) {
|
||||
const auto Size = GetSrcSize(Op);
|
||||
OrderedNode* Mem = MakeSegmentAddress(Op, Op->Src[0]);
|
||||
Ref Mem = MakeSegmentAddress(Op, Op->Src[0]);
|
||||
|
||||
auto NewFCW = _LoadMem(GPRClass, 2, Mem, 2);
|
||||
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
|
||||
|
||||
OrderedNode* MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 1));
|
||||
auto NewFSW = _LoadMem(GPRClass, Size, MemLocation, Size);
|
||||
auto NewFSW = _LoadMem(GPRClass, Size, Mem, _Constant(Size * 1), Size, MEM_OFFSET_SXTX, 1);
|
||||
ReconstructX87StateFromFSW(NewFSW);
|
||||
|
||||
{
|
||||
// FTW
|
||||
OrderedNode* MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 2));
|
||||
SetX87FTW(_LoadMem(GPRClass, Size, MemLocation, Size));
|
||||
SetX87FTW(_LoadMem(GPRClass, Size, Mem, _Constant(Size * 2), Size, MEM_OFFSET_SXTX, 1));
|
||||
}
|
||||
}
|
||||
|
||||
@@ -951,53 +953,45 @@ void OpDispatchBuilder::X87FNSTENV(OpcodeArgs) {
|
||||
// 4 bytes : data pointer selector
|
||||
|
||||
const auto Size = GetDstSize(Op);
|
||||
OrderedNode* Mem = MakeSegmentAddress(Op, Op->Dest);
|
||||
Ref Mem = MakeSegmentAddress(Op, Op->Dest);
|
||||
|
||||
{
|
||||
auto FCW = _LoadContext(2, GPRClass, offsetof(FEXCore::Core::CPUState, FCW));
|
||||
_StoreMem(GPRClass, Size, Mem, FCW, Size);
|
||||
}
|
||||
|
||||
{
|
||||
OrderedNode* MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 1));
|
||||
_StoreMem(GPRClass, Size, MemLocation, ReconstructFSW(), Size);
|
||||
}
|
||||
{ _StoreMem(GPRClass, Size, ReconstructFSW(), Mem, _Constant(Size * 1), Size, MEM_OFFSET_SXTX, 1); }
|
||||
|
||||
auto ZeroConst = _Constant(0);
|
||||
|
||||
{
|
||||
// FTW
|
||||
OrderedNode* MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 2));
|
||||
_StoreMem(GPRClass, Size, MemLocation, GetX87FTW(), Size);
|
||||
_StoreMem(GPRClass, Size, GetX87FTW(), Mem, _Constant(Size * 2), Size, MEM_OFFSET_SXTX, 1);
|
||||
}
|
||||
|
||||
{
|
||||
// Instruction Offset
|
||||
OrderedNode* MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 3));
|
||||
_StoreMem(GPRClass, Size, MemLocation, ZeroConst, Size);
|
||||
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(Size * 3), Size, MEM_OFFSET_SXTX, 1);
|
||||
}
|
||||
|
||||
{
|
||||
// Instruction CS selector (+ Opcode)
|
||||
OrderedNode* MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 4));
|
||||
_StoreMem(GPRClass, Size, MemLocation, ZeroConst, Size);
|
||||
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(Size * 4), Size, MEM_OFFSET_SXTX, 1);
|
||||
}
|
||||
|
||||
{
|
||||
// Data pointer offset
|
||||
OrderedNode* MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 5));
|
||||
_StoreMem(GPRClass, Size, MemLocation, ZeroConst, Size);
|
||||
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(Size * 5), Size, MEM_OFFSET_SXTX, 1);
|
||||
}
|
||||
|
||||
{
|
||||
// Data pointer selector
|
||||
OrderedNode* MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 6));
|
||||
_StoreMem(GPRClass, Size, MemLocation, ZeroConst, Size);
|
||||
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(Size * 6), Size, MEM_OFFSET_SXTX, 1);
|
||||
}
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::X87FLDCW(OpcodeArgs) {
|
||||
OrderedNode* NewFCW = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
|
||||
Ref NewFCW = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
|
||||
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
|
||||
}
|
||||
|
||||
@@ -1008,7 +1002,7 @@ void OpDispatchBuilder::X87FSTCW(OpcodeArgs) {
|
||||
}
|
||||
|
||||
void OpDispatchBuilder::X87LDSW(OpcodeArgs) {
|
||||
OrderedNode* NewFSW = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
|
||||
Ref NewFSW = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
|
||||
ReconstructX87StateFromFSW(NewFSW);
|
||||
}
|
||||
|
||||
@@ -1037,59 +1031,47 @@ void OpDispatchBuilder::X87FNSAVE(OpcodeArgs) {
|
||||
// 4 bytes : data pointer selector
|
||||
|
||||
const auto Size = GetDstSize(Op);
|
||||
OrderedNode* Mem = MakeSegmentAddress(Op, Op->Dest);
|
||||
OrderedNode* Top = GetX87Top();
|
||||
Ref Mem = MakeSegmentAddress(Op, Op->Dest);
|
||||
Ref Top = GetX87Top();
|
||||
{
|
||||
auto FCW = _LoadContext(2, GPRClass, offsetof(FEXCore::Core::CPUState, FCW));
|
||||
_StoreMem(GPRClass, Size, Mem, FCW, Size);
|
||||
}
|
||||
|
||||
{
|
||||
OrderedNode* MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 1));
|
||||
_StoreMem(GPRClass, Size, MemLocation, ReconstructFSW(), Size);
|
||||
}
|
||||
{ _StoreMem(GPRClass, Size, ReconstructFSW(), Mem, _Constant(Size * 1), Size, MEM_OFFSET_SXTX, 1); }
|
||||
|
||||
auto ZeroConst = _Constant(0);
|
||||
|
||||
{
|
||||
// FTW
|
||||
OrderedNode* MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 2));
|
||||
_StoreMem(GPRClass, Size, MemLocation, GetX87FTW(), Size);
|
||||
_StoreMem(GPRClass, Size, GetX87FTW(), Mem, _Constant(Size * 2), Size, MEM_OFFSET_SXTX, 1);
|
||||
}
|
||||
|
||||
{
|
||||
// Instruction Offset
|
||||
OrderedNode* MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 3));
|
||||
_StoreMem(GPRClass, Size, MemLocation, ZeroConst, Size);
|
||||
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(Size * 3), Size, MEM_OFFSET_SXTX, 1);
|
||||
}
|
||||
|
||||
{
|
||||
// Instruction CS selector (+ Opcode)
|
||||
OrderedNode* MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 4));
|
||||
_StoreMem(GPRClass, Size, MemLocation, ZeroConst, Size);
|
||||
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(Size * 4), Size, MEM_OFFSET_SXTX, 1);
|
||||
}
|
||||
|
||||
{
|
||||
// Data pointer offset
|
||||
OrderedNode* MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 5));
|
||||
_StoreMem(GPRClass, Size, MemLocation, ZeroConst, Size);
|
||||
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(Size * 5), Size, MEM_OFFSET_SXTX, 1);
|
||||
}
|
||||
|
||||
{
|
||||
// Data pointer selector
|
||||
OrderedNode* MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 6));
|
||||
_StoreMem(GPRClass, Size, MemLocation, ZeroConst, Size);
|
||||
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(Size * 6), Size, MEM_OFFSET_SXTX, 1);
|
||||
}
|
||||
|
||||
OrderedNode* ST0Location = _Add(OpSize::i64Bit, Mem, _Constant(Size * 7));
|
||||
|
||||
auto OneConst = _Constant(1);
|
||||
auto SevenConst = _Constant(7);
|
||||
auto TenConst = _Constant(10);
|
||||
for (int i = 0; i < 7; ++i) {
|
||||
auto data = _LoadContextIndexed(Top, 16, MMBaseOffset(), 16, FPRClass);
|
||||
_StoreMem(FPRClass, 16, ST0Location, data, 1);
|
||||
ST0Location = _Add(OpSize::i64Bit, ST0Location, TenConst);
|
||||
_StoreMem(FPRClass, 16, data, Mem, _Constant((Size * 7) + (10 * i)), 1, MEM_OFFSET_SXTX, 1);
|
||||
Top = _And(OpSize::i32Bit, _Add(OpSize::i32Bit, Top, OneConst), SevenConst);
|
||||
}
|
||||
|
||||
@@ -1098,10 +1080,9 @@ void OpDispatchBuilder::X87FNSAVE(OpcodeArgs) {
|
||||
// ST7 broken in to two parts
|
||||
// Lower 64bits [63:0]
|
||||
// upper 16 bits [79:64]
|
||||
_StoreMem(FPRClass, 8, ST0Location, data, 1);
|
||||
ST0Location = _Add(OpSize::i64Bit, ST0Location, _Constant(8));
|
||||
_StoreMem(FPRClass, 8, data, Mem, _Constant((Size * 7) + (7 * 10)), 1, MEM_OFFSET_SXTX, 1);
|
||||
auto topBytes = _VDupElement(16, 2, data, 4);
|
||||
_StoreMem(FPRClass, 2, ST0Location, topBytes, 1);
|
||||
_StoreMem(FPRClass, 2, topBytes, Mem, _Constant((Size * 7) + (7 * 10) + 8), 1, MEM_OFFSET_SXTX, 1);
|
||||
|
||||
// reset to default
|
||||
FNINIT(Op);
|
||||
@@ -1109,39 +1090,33 @@ void OpDispatchBuilder::X87FNSAVE(OpcodeArgs) {
|
||||
|
||||
void OpDispatchBuilder::X87FRSTOR(OpcodeArgs) {
|
||||
const auto Size = GetSrcSize(Op);
|
||||
OrderedNode* Mem = MakeSegmentAddress(Op, Op->Src[0]);
|
||||
Ref Mem = MakeSegmentAddress(Op, Op->Src[0]);
|
||||
|
||||
auto NewFCW = _LoadMem(GPRClass, 2, Mem, 2);
|
||||
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
|
||||
|
||||
OrderedNode* MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 1));
|
||||
auto NewFSW = _LoadMem(GPRClass, Size, MemLocation, Size);
|
||||
auto NewFSW = _LoadMem(GPRClass, Size, Mem, _Constant(Size * 1), Size, MEM_OFFSET_SXTX, 1);
|
||||
auto Top = ReconstructX87StateFromFSW(NewFSW);
|
||||
{
|
||||
// FTW
|
||||
OrderedNode* MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 2));
|
||||
SetX87FTW(_LoadMem(GPRClass, Size, MemLocation, Size));
|
||||
SetX87FTW(_LoadMem(GPRClass, Size, Mem, _Constant(Size * 2), Size, MEM_OFFSET_SXTX, 1));
|
||||
}
|
||||
|
||||
OrderedNode* ST0Location = _Add(OpSize::i64Bit, Mem, _Constant(Size * 7));
|
||||
|
||||
auto OneConst = _Constant(1);
|
||||
auto SevenConst = _Constant(7);
|
||||
auto TenConst = _Constant(10);
|
||||
|
||||
auto low = _Constant(~0ULL);
|
||||
auto high = _Constant(0xFFFF);
|
||||
OrderedNode* Mask = _VCastFromGPR(16, 8, low);
|
||||
Ref Mask = _VCastFromGPR(16, 8, low);
|
||||
Mask = _VInsGPR(16, 8, 1, Mask, high);
|
||||
|
||||
for (int i = 0; i < 7; ++i) {
|
||||
OrderedNode* Reg = _LoadMem(FPRClass, 16, ST0Location, 1);
|
||||
Ref Reg = _LoadMem(FPRClass, 16, Mem, _Constant((Size * 7) + (10 * i)), 1, MEM_OFFSET_SXTX, 1);
|
||||
// Mask off the top bits
|
||||
Reg = _VAnd(16, 16, Reg, Mask);
|
||||
|
||||
_StoreContextIndexed(Reg, Top, 16, MMBaseOffset(), 16, FPRClass);
|
||||
|
||||
ST0Location = _Add(OpSize::i64Bit, ST0Location, TenConst);
|
||||
Top = _And(OpSize::i32Bit, _Add(OpSize::i32Bit, Top, OneConst), SevenConst);
|
||||
}
|
||||
|
||||
@@ -1150,9 +1125,8 @@ void OpDispatchBuilder::X87FRSTOR(OpcodeArgs) {
|
||||
// Lower 64bits [63:0]
|
||||
// upper 16 bits [79:64]
|
||||
|
||||
OrderedNode* Reg = _LoadMem(FPRClass, 8, ST0Location, 1);
|
||||
ST0Location = _Add(OpSize::i64Bit, ST0Location, _Constant(8));
|
||||
OrderedNode* RegHigh = _LoadMem(FPRClass, 2, ST0Location, 1);
|
||||
Ref Reg = _LoadMem(FPRClass, 8, Mem, _Constant((Size * 7) + (10 * 7)), 1, MEM_OFFSET_SXTX, 1);
|
||||
Ref RegHigh = _LoadMem(FPRClass, 2, Mem, _Constant((Size * 7) + (10 * 7) + 8), 1, MEM_OFFSET_SXTX, 1);
|
||||
Reg = _VInsElement(16, 2, 4, 0, Reg, RegHigh);
|
||||
_StoreContextIndexed(Reg, Top, 16, MMBaseOffset(), 16, FPRClass);
|
||||
}
|
||||
@@ -1160,7 +1134,7 @@ void OpDispatchBuilder::X87FRSTOR(OpcodeArgs) {
|
||||
void OpDispatchBuilder::X87FXAM(OpcodeArgs) {
|
||||
auto top = GetX87Top();
|
||||
auto a = _LoadContextIndexed(top, 16, MMBaseOffset(), 16, FPRClass);
|
||||
OrderedNode* Result = _VExtractToGPR(16, 8, a, 1);
|
||||
Ref Result = _VExtractToGPR(16, 8, a, 1);
|
||||
|
||||
// Extract the sign bit
|
||||
Result = _Bfe(OpSize::i64Bit, 1, 15, Result);
|
||||
@@ -1219,11 +1193,11 @@ void OpDispatchBuilder::X87FCMOV(OpcodeArgs) {
|
||||
auto ZeroConst = _Constant(0);
|
||||
auto AllOneConst = _Constant(0xffff'ffff'ffff'ffffull);
|
||||
|
||||
OrderedNode* SrcCond = SelectCC(CC, OpSize::i64Bit, AllOneConst, ZeroConst);
|
||||
OrderedNode* VecCond = _VDupFromGPR(16, 8, SrcCond);
|
||||
Ref SrcCond = SelectCC(CC, OpSize::i64Bit, AllOneConst, ZeroConst);
|
||||
Ref VecCond = _VDupFromGPR(16, 8, SrcCond);
|
||||
|
||||
auto top = GetX87Top();
|
||||
OrderedNode* arg;
|
||||
Ref arg;
|
||||
|
||||
auto mask = _Constant(7);
|
||||
|
||||
@@ -1246,7 +1220,7 @@ void OpDispatchBuilder::X87EMMS(OpcodeArgs) {
|
||||
|
||||
void OpDispatchBuilder::X87FFREE(OpcodeArgs) {
|
||||
// Only sets the selected stack register's tag bits to EMPTY
|
||||
OrderedNode* top = GetX87Top();
|
||||
Ref top = GetX87Top();
|
||||
|
||||
// Implicit arg
|
||||
auto offset = _Constant(Op->OP & 7);
|
||||
|
||||
@@ -65,32 +65,30 @@ void OpDispatchBuilder::FNINITF64(OpcodeArgs) {
|
||||
|
||||
void OpDispatchBuilder::X87LDENVF64(OpcodeArgs) {
|
||||
const auto Size = GetSrcSize(Op);
|
||||
OrderedNode* Mem = MakeSegmentAddress(Op, Op->Src[0]);
|
||||
Ref Mem = MakeSegmentAddress(Op, Op->Src[0]);
|
||||
|
||||
auto NewFCW = _LoadMem(GPRClass, 2, Mem, 2);
|
||||
// ignore the rounding precision, we're always 64-bit in F64.
|
||||
// extract rounding mode
|
||||
OrderedNode* roundingMode = _Bfe(OpSize::i32Bit, 3, 10, NewFCW);
|
||||
Ref roundingMode = _Bfe(OpSize::i32Bit, 3, 10, NewFCW);
|
||||
_SetRoundingMode(roundingMode);
|
||||
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
|
||||
|
||||
OrderedNode* MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 1));
|
||||
auto NewFSW = _LoadMem(GPRClass, Size, MemLocation, Size);
|
||||
auto NewFSW = _LoadMem(GPRClass, Size, Mem, _Constant(Size * 1), Size, MEM_OFFSET_SXTX, 1);
|
||||
ReconstructX87StateFromFSW(NewFSW);
|
||||
|
||||
{
|
||||
// FTW
|
||||
OrderedNode* MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 2));
|
||||
SetX87FTW(_LoadMem(GPRClass, Size, MemLocation, Size));
|
||||
SetX87FTW(_LoadMem(GPRClass, Size, Mem, _Constant(Size * 2), Size, MEM_OFFSET_SXTX, 1));
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
void OpDispatchBuilder::X87FLDCWF64(OpcodeArgs) {
|
||||
OrderedNode* NewFCW = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
|
||||
Ref NewFCW = LoadSource(GPRClass, Op, Op->Src[0], Op->Flags);
|
||||
// ignore the rounding precision, we're always 64-bit in F64.
|
||||
// extract rounding mode
|
||||
OrderedNode* roundingMode = _Bfe(OpSize::i32Bit, 3, 10, NewFCW);
|
||||
Ref roundingMode = _Bfe(OpSize::i32Bit, 3, 10, NewFCW);
|
||||
_SetRoundingMode(roundingMode);
|
||||
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
|
||||
}
|
||||
@@ -105,8 +103,8 @@ void OpDispatchBuilder::FLDF64(OpcodeArgs) {
|
||||
|
||||
size_t read_width = (width == 80) ? 16 : width / 8;
|
||||
|
||||
OrderedNode* data {};
|
||||
OrderedNode* converted {};
|
||||
Ref data {};
|
||||
Ref converted {};
|
||||
|
||||
if (!Op->Src[0].IsNone()) {
|
||||
// Read from memory
|
||||
@@ -147,8 +145,8 @@ void OpDispatchBuilder::FBLDF64(OpcodeArgs) {
|
||||
SetX87Top(top);
|
||||
|
||||
// Read from memory
|
||||
OrderedNode* data = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], 16, Op->Flags);
|
||||
OrderedNode* converted = _F80BCDLoad(data);
|
||||
Ref data = LoadSource_WithOpSize(FPRClass, Op, Op->Src[0], 16, Op->Flags);
|
||||
Ref converted = _F80BCDLoad(data);
|
||||
converted = _F80CVT(8, converted);
|
||||
_StoreContextIndexed(converted, top, 8, MMBaseOffset(), 16, FPRClass);
|
||||
}
|
||||
@@ -157,7 +155,7 @@ void OpDispatchBuilder::FBSTPF64(OpcodeArgs) {
|
||||
auto orig_top = GetX87Top();
|
||||
auto data = _LoadContextIndexed(orig_top, 8, MMBaseOffset(), 16, FPRClass);
|
||||
|
||||
OrderedNode* converted = _F80CVTTo(data, 8);
|
||||
Ref converted = _F80CVTTo(data, 8);
|
||||
converted = _F80BCDStore(converted);
|
||||
|
||||
StoreResult_WithOpSize(FPRClass, Op, Op->Dest, converted, 10, 1);
|
||||
@@ -241,7 +239,7 @@ void OpDispatchBuilder::FISTF64(OpcodeArgs) {
|
||||
auto Size = GetSrcSize(Op);
|
||||
|
||||
auto orig_top = GetX87Top();
|
||||
OrderedNode* data = _LoadContextIndexed(orig_top, 8, MMBaseOffset(), 16, FPRClass);
|
||||
Ref data = _LoadContextIndexed(orig_top, 8, MMBaseOffset(), 16, FPRClass);
|
||||
if constexpr (Truncate) {
|
||||
data = _Float_ToGPR_ZS(Size == 4 ? 4 : 8, 8, data);
|
||||
} else {
|
||||
@@ -264,10 +262,10 @@ template void OpDispatchBuilder::FISTF64<true>(OpcodeArgs);
|
||||
template<size_t width, bool Integer, OpDispatchBuilder::OpResult ResInST0>
|
||||
void OpDispatchBuilder::FADDF64(OpcodeArgs) {
|
||||
auto top = GetX87Top();
|
||||
OrderedNode* StackLocation = top;
|
||||
Ref StackLocation = top;
|
||||
|
||||
OrderedNode* arg {};
|
||||
OrderedNode* b {};
|
||||
Ref arg {};
|
||||
Ref b {};
|
||||
|
||||
auto mask = _Constant(7);
|
||||
|
||||
@@ -320,9 +318,9 @@ template void OpDispatchBuilder::FADDF64<32, true, OpDispatchBuilder::OpResult::
|
||||
template<size_t width, bool Integer, OpDispatchBuilder::OpResult ResInST0>
|
||||
void OpDispatchBuilder::FMULF64(OpcodeArgs) {
|
||||
auto top = GetX87Top();
|
||||
OrderedNode* StackLocation = top;
|
||||
OrderedNode* arg {};
|
||||
OrderedNode* b {};
|
||||
Ref StackLocation = top;
|
||||
Ref arg {};
|
||||
Ref b {};
|
||||
|
||||
auto mask = _Constant(7);
|
||||
|
||||
@@ -378,9 +376,9 @@ template void OpDispatchBuilder::FMULF64<32, true, OpDispatchBuilder::OpResult::
|
||||
template<size_t width, bool Integer, bool reverse, OpDispatchBuilder::OpResult ResInST0>
|
||||
void OpDispatchBuilder::FDIVF64(OpcodeArgs) {
|
||||
auto top = GetX87Top();
|
||||
OrderedNode* StackLocation = top;
|
||||
OrderedNode* arg {};
|
||||
OrderedNode* b {};
|
||||
Ref StackLocation = top;
|
||||
Ref arg {};
|
||||
Ref b {};
|
||||
|
||||
auto mask = _Constant(7);
|
||||
|
||||
@@ -411,7 +409,7 @@ void OpDispatchBuilder::FDIVF64(OpcodeArgs) {
|
||||
|
||||
auto a = _LoadContextIndexed(top, 8, MMBaseOffset(), 16, FPRClass);
|
||||
|
||||
OrderedNode* result {};
|
||||
Ref result {};
|
||||
if constexpr (reverse) {
|
||||
result = _VFDiv(8, 8, b, a);
|
||||
} else {
|
||||
@@ -451,9 +449,9 @@ template void OpDispatchBuilder::FDIVF64<32, true, true, OpDispatchBuilder::OpRe
|
||||
template<size_t width, bool Integer, bool reverse, OpDispatchBuilder::OpResult ResInST0>
|
||||
void OpDispatchBuilder::FSUBF64(OpcodeArgs) {
|
||||
auto top = GetX87Top();
|
||||
OrderedNode* StackLocation = top;
|
||||
OrderedNode* arg {};
|
||||
OrderedNode* b {};
|
||||
Ref StackLocation = top;
|
||||
Ref arg {};
|
||||
Ref b {};
|
||||
|
||||
auto mask = _Constant(7);
|
||||
|
||||
@@ -484,7 +482,7 @@ void OpDispatchBuilder::FSUBF64(OpcodeArgs) {
|
||||
|
||||
auto a = _LoadContextIndexed(top, 8, MMBaseOffset(), 16, FPRClass);
|
||||
|
||||
OrderedNode* result {};
|
||||
Ref result {};
|
||||
if constexpr (reverse) {
|
||||
result = _VFSub(8, 8, b, a);
|
||||
} else {
|
||||
@@ -543,7 +541,7 @@ void OpDispatchBuilder::FTSTF64(OpcodeArgs) {
|
||||
auto a = _LoadContextIndexed(top, 8, MMBaseOffset(), 16, FPRClass);
|
||||
|
||||
auto low = _Constant(0);
|
||||
OrderedNode* data = _VCastFromGPR(8, 8, low);
|
||||
Ref data = _VCastFromGPR(8, 8, low);
|
||||
|
||||
// We are going to clobber NZCV, make sure it's in a GPR first.
|
||||
GetNZCV();
|
||||
@@ -573,11 +571,11 @@ void OpDispatchBuilder::FXTRACTF64(OpcodeArgs) {
|
||||
|
||||
auto a = _LoadContextIndexed(orig_top, 8, MMBaseOffset(), 16, FPRClass);
|
||||
auto gpr = _VExtractToGPR(8, 8, a, 0);
|
||||
OrderedNode* exp = _And(OpSize::i64Bit, gpr, _Constant(0x7ff0000000000000LL));
|
||||
Ref exp = _And(OpSize::i64Bit, gpr, _Constant(0x7ff0000000000000LL));
|
||||
exp = _Lshr(OpSize::i64Bit, exp, _Constant(52));
|
||||
exp = _Sub(OpSize::i64Bit, exp, _Constant(1023));
|
||||
exp = _Float_FromGPR_S(8, 8, exp);
|
||||
OrderedNode* sig = _And(OpSize::i64Bit, gpr, _Constant(0x800fffffffffffffLL));
|
||||
Ref sig = _And(OpSize::i64Bit, gpr, _Constant(0x800fffffffffffffLL));
|
||||
sig = _Or(OpSize::i64Bit, sig, _Constant(0x3ff0000000000000LL));
|
||||
sig = _VCastFromGPR(8, 8, sig);
|
||||
// Write to ST[TOP]
|
||||
@@ -591,8 +589,8 @@ void OpDispatchBuilder::FCOMIF64(OpcodeArgs) {
|
||||
auto top = GetX87Top();
|
||||
auto mask = _Constant(7);
|
||||
|
||||
OrderedNode* arg {};
|
||||
OrderedNode* b {};
|
||||
Ref arg {};
|
||||
Ref b {};
|
||||
|
||||
if (!Op->Src[0].IsNone()) {
|
||||
// Memory arg
|
||||
@@ -625,13 +623,7 @@ void OpDispatchBuilder::FCOMIF64(OpcodeArgs) {
|
||||
PossiblySetNZCVBits = ~0;
|
||||
ConvertNZCVToX87();
|
||||
} else {
|
||||
// Invalidate deferred flags early
|
||||
// OF, SF, AF, PF all undefined
|
||||
InvalidateDeferredFlags();
|
||||
|
||||
_FCmp(8, a, b);
|
||||
PossiblySetNZCVBits = ~0;
|
||||
ConvertNZCVToSSE();
|
||||
Comiss(8, a, b, true /* InvalidateAF */);
|
||||
}
|
||||
|
||||
if constexpr (poptwice) {
|
||||
@@ -701,7 +693,7 @@ void OpDispatchBuilder::X87BinaryOpF64(OpcodeArgs) {
|
||||
auto top = GetX87Top();
|
||||
|
||||
auto mask = _Constant(7);
|
||||
OrderedNode* st1 = _And(OpSize::i32Bit, _Add(OpSize::i32Bit, top, _Constant(1)), mask);
|
||||
Ref st1 = _And(OpSize::i32Bit, _Add(OpSize::i32Bit, top, _Constant(1)), mask);
|
||||
|
||||
auto a = _LoadContextIndexed(top, 8, MMBaseOffset(), 16, FPRClass);
|
||||
st1 = _LoadContextIndexed(st1, 8, MMBaseOffset(), 16, FPRClass);
|
||||
@@ -749,8 +741,8 @@ void OpDispatchBuilder::X87FYL2XF64(OpcodeArgs) {
|
||||
auto top = _And(OpSize::i32Bit, _Add(OpSize::i32Bit, orig_top, _Constant(1)), _Constant(7));
|
||||
SetX87Top(top);
|
||||
|
||||
OrderedNode* st0 = _LoadContextIndexed(orig_top, 8, MMBaseOffset(), 16, FPRClass);
|
||||
OrderedNode* st1 = _LoadContextIndexed(top, 8, MMBaseOffset(), 16, FPRClass);
|
||||
Ref st0 = _LoadContextIndexed(orig_top, 8, MMBaseOffset(), 16, FPRClass);
|
||||
Ref st1 = _LoadContextIndexed(top, 8, MMBaseOffset(), 16, FPRClass);
|
||||
|
||||
if (Plus1) {
|
||||
auto one = _VCastFromGPR(8, 8, _Constant(0x3FF0000000000000));
|
||||
@@ -791,7 +783,7 @@ void OpDispatchBuilder::X87ATANF64(OpcodeArgs) {
|
||||
SetX87Top(top);
|
||||
|
||||
auto a = _LoadContextIndexed(orig_top, 8, MMBaseOffset(), 16, FPRClass);
|
||||
OrderedNode* st1 = _LoadContextIndexed(top, 8, MMBaseOffset(), 16, FPRClass);
|
||||
Ref st1 = _LoadContextIndexed(top, 8, MMBaseOffset(), 16, FPRClass);
|
||||
|
||||
auto result = _F64ATAN(st1, a);
|
||||
|
||||
@@ -822,73 +814,60 @@ void OpDispatchBuilder::X87FNSAVEF64(OpcodeArgs) {
|
||||
// 4 bytes : data pointer selector
|
||||
|
||||
const auto Size = GetDstSize(Op);
|
||||
OrderedNode* Mem = MakeSegmentAddress(Op, Op->Dest);
|
||||
OrderedNode* Top = GetX87Top();
|
||||
Ref Mem = MakeSegmentAddress(Op, Op->Dest);
|
||||
Ref Top = GetX87Top();
|
||||
{
|
||||
auto FCW = _LoadContext(2, GPRClass, offsetof(FEXCore::Core::CPUState, FCW));
|
||||
_StoreMem(GPRClass, Size, Mem, FCW, Size);
|
||||
}
|
||||
|
||||
{
|
||||
OrderedNode* MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 1));
|
||||
_StoreMem(GPRClass, Size, MemLocation, ReconstructFSW(), Size);
|
||||
}
|
||||
{ _StoreMem(GPRClass, Size, ReconstructFSW(), Mem, _Constant(Size * 1), Size, MEM_OFFSET_SXTX, 1); }
|
||||
|
||||
auto ZeroConst = _Constant(0);
|
||||
|
||||
{
|
||||
// FTW
|
||||
OrderedNode* MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 2));
|
||||
_StoreMem(GPRClass, Size, MemLocation, GetX87FTW(), Size);
|
||||
_StoreMem(GPRClass, Size, GetX87FTW(), Mem, _Constant(Size * 2), Size, MEM_OFFSET_SXTX, 1);
|
||||
}
|
||||
|
||||
{
|
||||
// Instruction Offset
|
||||
OrderedNode* MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 3));
|
||||
_StoreMem(GPRClass, Size, MemLocation, ZeroConst, Size);
|
||||
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(Size * 3), Size, MEM_OFFSET_SXTX, 1);
|
||||
}
|
||||
|
||||
{
|
||||
// Instruction CS selector (+ Opcode)
|
||||
OrderedNode* MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 4));
|
||||
_StoreMem(GPRClass, Size, MemLocation, ZeroConst, Size);
|
||||
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(Size * 4), Size, MEM_OFFSET_SXTX, 1);
|
||||
}
|
||||
|
||||
{
|
||||
// Data pointer offset
|
||||
OrderedNode* MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 5));
|
||||
_StoreMem(GPRClass, Size, MemLocation, ZeroConst, Size);
|
||||
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(Size * 5), Size, MEM_OFFSET_SXTX, 1);
|
||||
}
|
||||
|
||||
{
|
||||
// Data pointer selector
|
||||
OrderedNode* MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 6));
|
||||
_StoreMem(GPRClass, Size, MemLocation, ZeroConst, Size);
|
||||
_StoreMem(GPRClass, Size, ZeroConst, Mem, _Constant(Size * 6), Size, MEM_OFFSET_SXTX, 1);
|
||||
}
|
||||
|
||||
OrderedNode* ST0Location = _Add(OpSize::i64Bit, Mem, _Constant(Size * 7));
|
||||
|
||||
auto OneConst = _Constant(1);
|
||||
auto SevenConst = _Constant(7);
|
||||
auto TenConst = _Constant(10);
|
||||
for (int i = 0; i < 7; ++i) {
|
||||
OrderedNode* data = _LoadContextIndexed(Top, 8, MMBaseOffset(), 16, FPRClass);
|
||||
Ref data = _LoadContextIndexed(Top, 8, MMBaseOffset(), 16, FPRClass);
|
||||
data = _F80CVTTo(data, 8);
|
||||
_StoreMem(FPRClass, 16, ST0Location, data, 1);
|
||||
ST0Location = _Add(OpSize::i64Bit, ST0Location, TenConst);
|
||||
_StoreMem(FPRClass, 16, data, Mem, _Constant((Size * 7) + (i * 10)), 1, MEM_OFFSET_SXTX, 1);
|
||||
Top = _And(OpSize::i32Bit, _Add(OpSize::i32Bit, Top, OneConst), SevenConst);
|
||||
}
|
||||
|
||||
// The final st(7) needs a bit of special handling here
|
||||
OrderedNode* data = _LoadContextIndexed(Top, 8, MMBaseOffset(), 16, FPRClass);
|
||||
Ref data = _LoadContextIndexed(Top, 8, MMBaseOffset(), 16, FPRClass);
|
||||
data = _F80CVTTo(data, 8);
|
||||
// ST7 broken in to two parts
|
||||
// Lower 64bits [63:0]
|
||||
// upper 16 bits [79:64]
|
||||
_StoreMem(FPRClass, 8, ST0Location, data, 1);
|
||||
ST0Location = _Add(OpSize::i64Bit, ST0Location, _Constant(8));
|
||||
_StoreMem(FPRClass, 8, data, Mem, _Constant((Size * 7) + (7 * 10)), 1, MEM_OFFSET_SXTX, 1);
|
||||
auto topBytes = _VDupElement(16, 2, data, 4);
|
||||
_StoreMem(FPRClass, 2, ST0Location, topBytes, 1);
|
||||
_StoreMem(FPRClass, 2, topBytes, Mem, _Constant((Size * 7) + (7 * 10) + 8), 1, MEM_OFFSET_SXTX, 1);
|
||||
|
||||
// reset to default
|
||||
FNINIT(Op);
|
||||
@@ -898,12 +877,12 @@ void OpDispatchBuilder::X87FNSAVEF64(OpcodeArgs) {
|
||||
|
||||
void OpDispatchBuilder::X87FRSTORF64(OpcodeArgs) {
|
||||
const auto Size = GetSrcSize(Op);
|
||||
OrderedNode* Mem = MakeSegmentAddress(Op, Op->Src[0]);
|
||||
Ref Mem = MakeSegmentAddress(Op, Op->Src[0]);
|
||||
|
||||
auto NewFCW = _LoadMem(GPRClass, 2, Mem, 2);
|
||||
// ignore the rounding precision, we're always 64-bit in F64.
|
||||
// extract rounding mode
|
||||
OrderedNode* roundingMode = NewFCW;
|
||||
Ref roundingMode = NewFCW;
|
||||
auto roundShift = _Constant(10);
|
||||
auto roundMask = _Constant(3);
|
||||
roundingMode = _Lshr(OpSize::i32Bit, roundingMode, roundShift);
|
||||
@@ -912,36 +891,30 @@ void OpDispatchBuilder::X87FRSTORF64(OpcodeArgs) {
|
||||
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
|
||||
_StoreContext(2, GPRClass, NewFCW, offsetof(FEXCore::Core::CPUState, FCW));
|
||||
|
||||
OrderedNode* MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 1));
|
||||
auto NewFSW = _LoadMem(GPRClass, Size, MemLocation, Size);
|
||||
auto NewFSW = _LoadMem(GPRClass, Size, Mem, _Constant(Size * 1), Size, MEM_OFFSET_SXTX, 1);
|
||||
auto Top = ReconstructX87StateFromFSW(NewFSW);
|
||||
|
||||
{
|
||||
// FTW
|
||||
OrderedNode* MemLocation = _Add(OpSize::i64Bit, Mem, _Constant(Size * 2));
|
||||
SetX87FTW(_LoadMem(GPRClass, Size, MemLocation, Size));
|
||||
SetX87FTW(_LoadMem(GPRClass, Size, Mem, _Constant(Size * 2), Size, MEM_OFFSET_SXTX, 1));
|
||||
}
|
||||
|
||||
OrderedNode* ST0Location = _Add(OpSize::i64Bit, Mem, _Constant(Size * 7));
|
||||
|
||||
auto OneConst = _Constant(1);
|
||||
auto SevenConst = _Constant(7);
|
||||
auto TenConst = _Constant(10);
|
||||
|
||||
auto low = _Constant(~0ULL);
|
||||
auto high = _Constant(0xFFFF);
|
||||
OrderedNode* Mask = _VCastFromGPR(16, 8, low);
|
||||
Ref Mask = _VCastFromGPR(16, 8, low);
|
||||
Mask = _VInsGPR(16, 8, 1, Mask, high);
|
||||
|
||||
for (int i = 0; i < 7; ++i) {
|
||||
OrderedNode* Reg = _LoadMem(FPRClass, 16, ST0Location, 1);
|
||||
Ref Reg = _LoadMem(FPRClass, 16, Mem, _Constant((Size * 7) + (i * 10)), 1, MEM_OFFSET_SXTX, 1);
|
||||
// Mask off the top bits
|
||||
Reg = _VAnd(16, 16, Reg, Mask);
|
||||
// Convert to double precision
|
||||
Reg = _F80CVT(8, Reg);
|
||||
_StoreContextIndexed(Reg, Top, 8, MMBaseOffset(), 16, FPRClass);
|
||||
|
||||
ST0Location = _Add(OpSize::i64Bit, ST0Location, TenConst);
|
||||
Top = _And(OpSize::i32Bit, _Add(OpSize::i32Bit, Top, OneConst), SevenConst);
|
||||
}
|
||||
|
||||
@@ -950,9 +923,8 @@ void OpDispatchBuilder::X87FRSTORF64(OpcodeArgs) {
|
||||
// Lower 64bits [63:0]
|
||||
// upper 16 bits [79:64]
|
||||
|
||||
OrderedNode* Reg = _LoadMem(FPRClass, 8, ST0Location, 1);
|
||||
ST0Location = _Add(OpSize::i64Bit, ST0Location, _Constant(8));
|
||||
OrderedNode* RegHigh = _LoadMem(FPRClass, 2, ST0Location, 1);
|
||||
Ref Reg = _LoadMem(FPRClass, 8, Mem, _Constant((Size * 7) + (7 * 10)), 1, MEM_OFFSET_SXTX, 1);
|
||||
Ref RegHigh = _LoadMem(FPRClass, 2, Mem, _Constant((Size * 7) + (7 * 10) + 8), 1, MEM_OFFSET_SXTX, 1);
|
||||
Reg = _VInsElement(16, 2, 4, 0, Reg, RegHigh);
|
||||
Reg = _F80CVT(8, Reg); // Convert to double precision
|
||||
_StoreContextIndexed(Reg, Top, 8, MMBaseOffset(), 16, FPRClass);
|
||||
@@ -963,7 +935,7 @@ void OpDispatchBuilder::X87FRSTORF64(OpcodeArgs) {
|
||||
void OpDispatchBuilder::X87FXAMF64(OpcodeArgs) {
|
||||
auto top = GetX87Top();
|
||||
auto a = _LoadContextIndexed(top, 8, MMBaseOffset(), 16, FPRClass);
|
||||
OrderedNode* Result = _VExtractToGPR(8, 8, a, 0);
|
||||
Ref Result = _VExtractToGPR(8, 8, a, 0);
|
||||
|
||||
// Extract the sign bit
|
||||
Result = _Bfe(OpSize::i64Bit, 1, 63, Result);
|
||||
|
||||
@@ -102,10 +102,10 @@ std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps = []() consteval {
|
||||
{0x6B, 1, X86InstInfo{"IMUL", TYPE_INST, FLAGS_MODRM | FLAGS_SRC_SEXT , 1, nullptr}},
|
||||
|
||||
// This should just throw a GP
|
||||
{0x6C, 1, X86InstInfo{"INSB", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
|
||||
{0x6D, 1, X86InstInfo{"INSW", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
|
||||
{0x6E, 1, X86InstInfo{"OUTS", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
|
||||
{0x6F, 1, X86InstInfo{"OUTS", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
|
||||
{0x6C, 1, X86InstInfo{"INSB", TYPE_INST, FLAGS_NONE, 0, nullptr}},
|
||||
{0x6D, 1, X86InstInfo{"INSW", TYPE_INST, FLAGS_NONE, 0, nullptr}},
|
||||
{0x6E, 1, X86InstInfo{"OUTS", TYPE_INST, FLAGS_NONE, 0, nullptr}},
|
||||
{0x6F, 1, X86InstInfo{"OUTS", TYPE_INST, FLAGS_NONE, 0, nullptr}},
|
||||
|
||||
{0x70, 1, X86InstInfo{"JO", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
|
||||
{0x71, 1, X86InstInfo{"JNO", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
|
||||
@@ -183,24 +183,24 @@ std::array<X86InstInfo, MAX_PRIMARY_TABLE_SIZE> BaseOps = []() consteval {
|
||||
{0xE3, 1, X86InstInfo{"JrCXZ", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT , 1, nullptr}},
|
||||
|
||||
// Should just throw GP
|
||||
{0xE4, 2, X86InstInfo{"IN", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
|
||||
{0xE6, 2, X86InstInfo{"OUT", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
|
||||
{0xE4, 2, X86InstInfo{"IN", TYPE_INST, FLAGS_NONE, 1, nullptr}},
|
||||
{0xE6, 2, X86InstInfo{"OUT", TYPE_INST, FLAGS_NONE, 1, nullptr}},
|
||||
|
||||
{0xE8, 1, X86InstInfo{"CALL", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_BLOCK_END , 4, nullptr}},
|
||||
{0xE9, 1, X86InstInfo{"JMP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_DISPLACE_SIZE_DIV_2 | FLAGS_BLOCK_END , 4, nullptr}},
|
||||
{0xEB, 1, X86InstInfo{"JMP", TYPE_INST, GenFlagsSameSize(SIZE_64BITDEF) | FLAGS_SETS_RIP | FLAGS_SRC_SEXT | FLAGS_BLOCK_END , 1, nullptr}},
|
||||
|
||||
// Should just throw GP
|
||||
{0xEC, 2, X86InstInfo{"IN", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
|
||||
{0xEE, 2, X86InstInfo{"OUT", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
|
||||
{0xEC, 2, X86InstInfo{"IN", TYPE_INST, FLAGS_NONE, 0, nullptr}},
|
||||
{0xEE, 2, X86InstInfo{"OUT", TYPE_INST, FLAGS_NONE, 0, nullptr}},
|
||||
|
||||
{0xF1, 1, X86InstInfo{"INT1", TYPE_INST, FLAGS_NONE, 0, nullptr}},
|
||||
{0xF4, 1, X86InstInfo{"HLT", TYPE_INST, FLAGS_BLOCK_END, 0, nullptr}},
|
||||
{0xF5, 1, X86InstInfo{"CMC", TYPE_INST, FLAGS_NONE, 0, nullptr}},
|
||||
{0xF8, 1, X86InstInfo{"CLC", TYPE_INST, FLAGS_NONE, 0, nullptr}},
|
||||
{0xF9, 1, X86InstInfo{"STC", TYPE_INST, FLAGS_NONE, 0, nullptr}},
|
||||
{0xFA, 1, X86InstInfo{"CLI", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
|
||||
{0xFB, 1, X86InstInfo{"STI", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
|
||||
{0xFA, 1, X86InstInfo{"CLI", TYPE_INST, FLAGS_NONE, 0, nullptr}},
|
||||
{0xFB, 1, X86InstInfo{"STI", TYPE_INST, FLAGS_NONE, 0, nullptr}},
|
||||
{0xFC, 1, X86InstInfo{"CLD", TYPE_INST, FLAGS_NONE, 0, nullptr}},
|
||||
{0xFD, 1, X86InstInfo{"STD", TYPE_INST, FLAGS_NONE, 0, nullptr}},
|
||||
|
||||
|
||||
@@ -32,7 +32,7 @@ std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = [
|
||||
{OPD(TYPE_GROUP_6, PF_NONE, 0), 1, X86InstInfo{"SLDT", TYPE_UNDEC, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_6, PF_NONE, 1), 1, X86InstInfo{"STR", TYPE_PRIV, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_6, PF_NONE, 2), 1, X86InstInfo{"LLDT", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_6, PF_NONE, 3), 1, X86InstInfo{"LTR", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_6, PF_NONE, 3), 1, X86InstInfo{"LTR", TYPE_INST, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_6, PF_NONE, 4), 1, X86InstInfo{"VERR", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_6, PF_NONE, 5), 1, X86InstInfo{"VERW", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_6, PF_NONE, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
|
||||
@@ -41,7 +41,7 @@ std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = [
|
||||
{OPD(TYPE_GROUP_6, PF_F3, 0), 1, X86InstInfo{"SLDT", TYPE_UNDEC, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_6, PF_F3, 1), 1, X86InstInfo{"STR", TYPE_PRIV, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_6, PF_F3, 2), 1, X86InstInfo{"LLDT", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_6, PF_F3, 3), 1, X86InstInfo{"LTR", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_6, PF_F3, 3), 1, X86InstInfo{"LTR", TYPE_INST, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_6, PF_F3, 4), 1, X86InstInfo{"VERR", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_6, PF_F3, 5), 1, X86InstInfo{"VERW", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_6, PF_F3, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
|
||||
@@ -50,7 +50,7 @@ std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = [
|
||||
{OPD(TYPE_GROUP_6, PF_66, 0), 1, X86InstInfo{"SLDT", TYPE_UNDEC, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_6, PF_66, 1), 1, X86InstInfo{"STR", TYPE_PRIV, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_6, PF_66, 2), 1, X86InstInfo{"LLDT", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_6, PF_66, 3), 1, X86InstInfo{"LTR", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_6, PF_66, 3), 1, X86InstInfo{"LTR", TYPE_INST, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_6, PF_66, 4), 1, X86InstInfo{"VERR", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_6, PF_66, 5), 1, X86InstInfo{"VERW", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_6, PF_66, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
|
||||
@@ -59,7 +59,7 @@ std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = [
|
||||
{OPD(TYPE_GROUP_6, PF_F2, 0), 1, X86InstInfo{"SLDT", TYPE_UNDEC, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_6, PF_F2, 1), 1, X86InstInfo{"STR", TYPE_PRIV, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_6, PF_F2, 2), 1, X86InstInfo{"LLDT", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_6, PF_F2, 3), 1, X86InstInfo{"LTR", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_6, PF_F2, 3), 1, X86InstInfo{"LTR", TYPE_INST, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_6, PF_F2, 4), 1, X86InstInfo{"VERR", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_6, PF_F2, 5), 1, X86InstInfo{"VERW", TYPE_UNDEC, FLAGS_MODRM, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_6, PF_F2, 6), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
|
||||
@@ -72,7 +72,7 @@ std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = [
|
||||
{OPD(TYPE_GROUP_7, PF_NONE, 3), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_7, PF_NONE, 4), 1, X86InstInfo{"SMSW", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_7, PF_NONE, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_7, PF_NONE, 6), 1, X86InstInfo{"LMSW", TYPE_PRIV, FLAGS_MODRM, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_7, PF_NONE, 6), 1, X86InstInfo{"LMSW", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_7, PF_NONE, 7), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
|
||||
|
||||
{OPD(TYPE_GROUP_7, PF_F3, 0), 1, X86InstInfo{"SGDT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
|
||||
@@ -81,7 +81,7 @@ std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = [
|
||||
{OPD(TYPE_GROUP_7, PF_F3, 3), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_7, PF_F3, 4), 1, X86InstInfo{"SMSW", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_7, PF_F3, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_7, PF_F3, 6), 1, X86InstInfo{"LMSW", TYPE_PRIV, FLAGS_MODRM, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_7, PF_F3, 6), 1, X86InstInfo{"LMSW", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_7, PF_F3, 7), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
|
||||
|
||||
{OPD(TYPE_GROUP_7, PF_66, 0), 1, X86InstInfo{"SGDT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
|
||||
@@ -90,7 +90,7 @@ std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = [
|
||||
{OPD(TYPE_GROUP_7, PF_66, 3), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_7, PF_66, 4), 1, X86InstInfo{"SMSW", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_7, PF_66, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_7, PF_66, 6), 1, X86InstInfo{"LMSW", TYPE_PRIV, FLAGS_MODRM, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_7, PF_66, 6), 1, X86InstInfo{"LMSW", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_7, PF_66, 7), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
|
||||
|
||||
{OPD(TYPE_GROUP_7, PF_F2, 0), 1, X86InstInfo{"SGDT", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
|
||||
@@ -99,7 +99,7 @@ std::array<X86InstInfo, MAX_INST_SECOND_GROUP_TABLE_SIZE> SecondInstGroupOps = [
|
||||
{OPD(TYPE_GROUP_7, PF_F2, 3), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_7, PF_F2, 4), 1, X86InstInfo{"SMSW", TYPE_INST, FLAGS_MODRM | FLAGS_SF_MOD_DST, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_7, PF_F2, 5), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_7, PF_F2, 6), 1, X86InstInfo{"LMSW", TYPE_PRIV, FLAGS_MODRM, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_7, PF_F2, 6), 1, X86InstInfo{"LMSW", TYPE_INST, FLAGS_MODRM, 0, nullptr}},
|
||||
{OPD(TYPE_GROUP_7, PF_F2, 7), 1, X86InstInfo{"", TYPE_SECOND_GROUP_MODRM, FLAGS_NONE, 0, nullptr}},
|
||||
|
||||
// GROUP 8
|
||||
|
||||
@@ -15,8 +15,8 @@ std::array<X86InstInfo, MAX_SECOND_MODRM_TABLE_SIZE> SecondModRMTableOps = []()
|
||||
std::array<X86InstInfo, MAX_SECOND_MODRM_TABLE_SIZE> Table{};
|
||||
constexpr U8U8InfoStruct SecondaryModRMExtensionOpTable[] = {
|
||||
// REG /1
|
||||
{((0 << 3) | 0), 1, X86InstInfo{"MONITOR", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
|
||||
{((0 << 3) | 1), 1, X86InstInfo{"MWAIT", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
|
||||
{((0 << 3) | 0), 1, X86InstInfo{"MONITOR", TYPE_INST, FLAGS_NONE, 0, nullptr}},
|
||||
{((0 << 3) | 1), 1, X86InstInfo{"MWAIT", TYPE_INST, FLAGS_NONE, 0, nullptr}},
|
||||
{((0 << 3) | 2), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
|
||||
{((0 << 3) | 3), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
|
||||
{((0 << 3) | 4), 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
|
||||
@@ -42,10 +42,10 @@ std::array<X86InstInfo, MAX_SECOND_MODRM_TABLE_SIZE> SecondModRMTableOps = []()
|
||||
{((2 << 3) | 4), 1, X86InstInfo{"STGI", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
|
||||
{((2 << 3) | 5), 1, X86InstInfo{"CLGI", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
|
||||
{((2 << 3) | 6), 1, X86InstInfo{"SKINIT", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
|
||||
{((2 << 3) | 7), 1, X86InstInfo{"INVLPGA", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
|
||||
{((2 << 3) | 7), 1, X86InstInfo{"INVLPGA", TYPE_INST, FLAGS_NONE, 0, nullptr}},
|
||||
|
||||
// REG /7
|
||||
{((3 << 3) | 0), 1, X86InstInfo{"SWAPGS", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
|
||||
{((3 << 3) | 0), 1, X86InstInfo{"SWAPGS", TYPE_INST, FLAGS_NONE, 0, nullptr}},
|
||||
{((3 << 3) | 1), 1, X86InstInfo{"RDTSCP", TYPE_INST, FLAGS_NONE, 0, nullptr}},
|
||||
{((3 << 3) | 2), 1, X86InstInfo{"MONITORX", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
|
||||
{((3 << 3) | 3), 1, X86InstInfo{"MWAITX", TYPE_PRIV, FLAGS_NONE, 0, nullptr}},
|
||||
|
||||
@@ -25,8 +25,8 @@ auto BaseOpsLambda = []() consteval {
|
||||
{0x03, 1, X86InstInfo{"LSL", TYPE_UNDEC, FLAGS_NO_OVERLAY, 0, nullptr}},
|
||||
{0x04, 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NO_OVERLAY, 0, nullptr}},
|
||||
{0x05, 1, X86InstInfo{"SYSCALL", TYPE_INST, DEFAULT_SYSCALL_FLAGS, 0, nullptr}},
|
||||
{0x06, 1, X86InstInfo{"CLTS", TYPE_PRIV, FLAGS_NO_OVERLAY, 0, nullptr}},
|
||||
{0x07, 1, X86InstInfo{"SYSRET", TYPE_PRIV, FLAGS_NO_OVERLAY, 0, nullptr}},
|
||||
{0x06, 1, X86InstInfo{"CLTS", TYPE_INST, FLAGS_NO_OVERLAY, 0, nullptr}},
|
||||
{0x07, 1, X86InstInfo{"SYSRET", TYPE_INST, FLAGS_NO_OVERLAY, 0, nullptr}},
|
||||
{0x08, 1, X86InstInfo{"INVD", TYPE_PRIV, FLAGS_NO_OVERLAY, 0, nullptr}},
|
||||
{0x09, 1, X86InstInfo{"WBINVD", TYPE_PRIV, FLAGS_NO_OVERLAY, 0, nullptr}},
|
||||
{0x0A, 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NO_OVERLAY, 0, nullptr}},
|
||||
@@ -47,8 +47,8 @@ auto BaseOpsLambda = []() consteval {
|
||||
{0x18, 1, X86InstInfo{"", TYPE_GROUP_16, FLAGS_NO_OVERLAY, 0, nullptr}},
|
||||
{0x19, 7, X86InstInfo{"NOP", TYPE_INST, FLAGS_MODRM | FLAGS_NO_OVERLAY, 0, nullptr}},
|
||||
|
||||
{0x20, 2, X86InstInfo{"MOV", TYPE_PRIV, GenFlagsSameSize(SIZE_64BIT) | FLAGS_NO_OVERLAY, 0, nullptr}},
|
||||
{0x22, 2, X86InstInfo{"MOV", TYPE_PRIV, GenFlagsSameSize(SIZE_64BIT) | FLAGS_NO_OVERLAY, 0, nullptr}},
|
||||
{0x20, 2, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_NO_OVERLAY, 0, nullptr}},
|
||||
{0x22, 2, X86InstInfo{"MOV", TYPE_INST, GenFlagsSameSize(SIZE_64BIT) | FLAGS_NO_OVERLAY, 0, nullptr}},
|
||||
{0x24, 4, X86InstInfo{"", TYPE_INVALID, FLAGS_NONE, 0, nullptr}},
|
||||
{0x28, 1, X86InstInfo{"MOVAPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{0x29, 1, X86InstInfo{"MOVAPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
@@ -59,12 +59,12 @@ auto BaseOpsLambda = []() consteval {
|
||||
{0x2E, 1, X86InstInfo{"UCOMISS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{0x2F, 1, X86InstInfo{"COMISS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
|
||||
{0x30, 1, X86InstInfo{"WRMSR", TYPE_PRIV, FLAGS_NO_OVERLAY, 0, nullptr}},
|
||||
{0x30, 1, X86InstInfo{"WRMSR", TYPE_INST, FLAGS_NO_OVERLAY, 0, nullptr}},
|
||||
{0x31, 1, X86InstInfo{"RDTSC", TYPE_INST, FLAGS_NO_OVERLAY, 0, nullptr}},
|
||||
{0x32, 1, X86InstInfo{"RDMSR", TYPE_PRIV, FLAGS_NO_OVERLAY, 0, nullptr}},
|
||||
{0x33, 1, X86InstInfo{"RDPMC", TYPE_PRIV, FLAGS_NO_OVERLAY, 0, nullptr}},
|
||||
{0x34, 1, X86InstInfo{"SYSENTER", TYPE_PRIV, FLAGS_NO_OVERLAY, 0, nullptr}},
|
||||
{0x35, 1, X86InstInfo{"SYSEXIT", TYPE_PRIV, FLAGS_NO_OVERLAY, 0, nullptr}},
|
||||
{0x32, 1, X86InstInfo{"RDMSR", TYPE_INST, FLAGS_NO_OVERLAY, 0, nullptr}},
|
||||
{0x33, 1, X86InstInfo{"RDPMC", TYPE_INST, FLAGS_NO_OVERLAY, 0, nullptr}},
|
||||
{0x34, 1, X86InstInfo{"SYSENTER", TYPE_INST, FLAGS_NO_OVERLAY, 0, nullptr}},
|
||||
{0x35, 1, X86InstInfo{"SYSEXIT", TYPE_INST, FLAGS_NO_OVERLAY, 0, nullptr}},
|
||||
{0x36, 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NO_OVERLAY, 0, nullptr}},
|
||||
{0x38, 1, X86InstInfo{"", TYPE_0F38_TABLE, FLAGS_NO_OVERLAY, 0, nullptr}},
|
||||
{0x39, 1, X86InstInfo{"", TYPE_INVALID, FLAGS_NO_OVERLAY, 0, nullptr}},
|
||||
|
||||
@@ -27,8 +27,8 @@ std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps = []() consteval {
|
||||
{OPD(1, 0b10, 0x11), 1, X86InstInfo{"VMOVSS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(1, 0b11, 0x11), 1, X86InstInfo{"VMOVSD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
|
||||
{OPD(1, 0b00, 0x12), 1, X86InstInfo{"VMOVLPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_XMM_FLAGS | FLAGS_VEX_1ST_SRC, 0, nullptr}},
|
||||
{OPD(1, 0b01, 0x12), 1, X86InstInfo{"VMOVLPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_XMM_FLAGS | FLAGS_VEX_1ST_SRC, 0, nullptr}},
|
||||
{OPD(1, 0b00, 0x12), 1, X86InstInfo{"VMOVLPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_VEX_1ST_SRC, 0, nullptr}},
|
||||
{OPD(1, 0b01, 0x12), 1, X86InstInfo{"VMOVLPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS | FLAGS_VEX_1ST_SRC, 0, nullptr}},
|
||||
{OPD(1, 0b10, 0x12), 1, X86InstInfo{"VMOVSLDUP", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(1, 0b11, 0x12), 1, X86InstInfo{"VMOVDDUP", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
|
||||
@@ -282,7 +282,7 @@ std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps = []() consteval {
|
||||
{OPD(2, 0b01, 0x0E), 1, X86InstInfo{"VTESTPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x0F), 1, X86InstInfo{"VTESTPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
|
||||
{OPD(2, 0b01, 0x13), 1, X86InstInfo{"VCVTPH2PS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x13), 1, X86InstInfo{"VCVTPH2PS", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_64BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x16), 1, X86InstInfo{"VPERMPS", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x17), 1, X86InstInfo{"VPTEST", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
|
||||
@@ -343,46 +343,46 @@ std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps = []() consteval {
|
||||
{OPD(2, 0b01, 0x8C), 1, X86InstInfo{"VPMASKMOV", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x8E), 1, X86InstInfo{"VPMASKMOV", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_SF_MOD_MEM_ONLY | FLAGS_SF_MOD_DST | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
|
||||
{OPD(2, 0b01, 0x90), 1, X86InstInfo{"VPGATHERDD/Q", TYPE_UNDEC, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_2ND_SRC | FLAGS_VEX_VSIB | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x91), 1, X86InstInfo{"VPGATHERQD/Q", TYPE_UNDEC, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_2ND_SRC | FLAGS_VEX_VSIB | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x92), 1, X86InstInfo{"VGATHERDPS/D", TYPE_UNDEC, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_2ND_SRC | FLAGS_VEX_VSIB | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x93), 1, X86InstInfo{"VGATHERQPS/D", TYPE_UNDEC, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_2ND_SRC | FLAGS_VEX_VSIB | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x90), 1, X86InstInfo{"VPGATHERDD/Q", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_2ND_SRC | FLAGS_VEX_VSIB | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x91), 1, X86InstInfo{"VPGATHERQD/Q", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_2ND_SRC | FLAGS_VEX_VSIB | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x92), 1, X86InstInfo{"VGATHERDPS/D", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_2ND_SRC | FLAGS_VEX_VSIB | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x93), 1, X86InstInfo{"VGATHERQPS/D", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_2ND_SRC | FLAGS_VEX_VSIB | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
|
||||
{OPD(2, 0b01, 0x96), 1, X86InstInfo{"VFMADDSUB132", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x97), 1, X86InstInfo{"VFMSUBADD132", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x96), 1, X86InstInfo{"VFMADDSUB132", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x97), 1, X86InstInfo{"VFMSUBADD132", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
|
||||
{OPD(2, 0b01, 0x98), 1, X86InstInfo{"VFMADD132", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x99), 1, X86InstInfo{"VFMADD132", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x9A), 1, X86InstInfo{"VFMSUB132", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x9B), 1, X86InstInfo{"VFMSUB132", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x9C), 1, X86InstInfo{"VFNMADD132", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x9D), 1, X86InstInfo{"VFNMADD132", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x9E), 1, X86InstInfo{"VFNMSUB132", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x9F), 1, X86InstInfo{"VFNMSUB132", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x98), 1, X86InstInfo{"VFMADD132", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x99), 1, X86InstInfo{"VFMADD132_S", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x9A), 1, X86InstInfo{"VFMSUB132", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x9B), 1, X86InstInfo{"VFMSUB132_S", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x9C), 1, X86InstInfo{"VFNMADD132", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x9D), 1, X86InstInfo{"VFNMADD132_S", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x9E), 1, X86InstInfo{"VFNMSUB132", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0x9F), 1, X86InstInfo{"VFNMSUB132_S", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
|
||||
{OPD(2, 0b01, 0xA8), 1, X86InstInfo{"VFMADD213", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xA9), 1, X86InstInfo{"VFMADD213", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xAA), 1, X86InstInfo{"VFMSUB213", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xAB), 1, X86InstInfo{"VFMSUB213", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xAC), 1, X86InstInfo{"VFNMADD213", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xAD), 1, X86InstInfo{"VFNMADD213", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xAE), 1, X86InstInfo{"VFNMSUB213", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xAF), 1, X86InstInfo{"VFNMSUB213", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xA8), 1, X86InstInfo{"VFMADD213", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xA9), 1, X86InstInfo{"VFMADD213_S", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xAA), 1, X86InstInfo{"VFMSUB213", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xAB), 1, X86InstInfo{"VFMSUB213_S", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xAC), 1, X86InstInfo{"VFNMADD213", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xAD), 1, X86InstInfo{"VFNMADD213_S", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xAE), 1, X86InstInfo{"VFNMSUB213", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xAF), 1, X86InstInfo{"VFNMSUB213_S", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
|
||||
{OPD(2, 0b01, 0xB8), 1, X86InstInfo{"VFMADD231", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xB9), 1, X86InstInfo{"VFMADD231", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xBA), 1, X86InstInfo{"VFMSUB231", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xBB), 1, X86InstInfo{"VFMSUB231", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xBC), 1, X86InstInfo{"VFNMADD231", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xBD), 1, X86InstInfo{"VFNMADD231", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xBE), 1, X86InstInfo{"VFNMSUB231", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xBF), 1, X86InstInfo{"VFNMSUB231", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xB8), 1, X86InstInfo{"VFMADD231", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xB9), 1, X86InstInfo{"VFMADD231_S", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xBA), 1, X86InstInfo{"VFMSUB231", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xBB), 1, X86InstInfo{"VFMSUB231_S", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xBC), 1, X86InstInfo{"VFNMADD231", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xBD), 1, X86InstInfo{"VFNMADD231_S", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xBE), 1, X86InstInfo{"VFNMSUB231", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xBF), 1, X86InstInfo{"VFNMSUB231_S", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
|
||||
{OPD(2, 0b01, 0xA6), 1, X86InstInfo{"VFMADDSUB213", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xA7), 1, X86InstInfo{"VFMSUBADD213", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xA6), 1, X86InstInfo{"VFMADDSUB213", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xA7), 1, X86InstInfo{"VFMSUBADD213", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
|
||||
{OPD(2, 0b01, 0xB6), 1, X86InstInfo{"VFMADDSUB231", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xB7), 1, X86InstInfo{"VFMSUBADD231", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xB6), 1, X86InstInfo{"VFMADDSUB231", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xB7), 1, X86InstInfo{"VFMSUBADD231", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
|
||||
{OPD(2, 0b01, 0xDB), 1, X86InstInfo{"VAESIMC", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
{OPD(2, 0b01, 0xDC), 1, X86InstInfo{"VAESENC", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 0, nullptr}},
|
||||
@@ -433,7 +433,7 @@ std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps = []() consteval {
|
||||
|
||||
{OPD(3, 0b01, 0x18), 1, X86InstInfo{"VINSERTF128", TYPE_INST, GenFlagsSameSize(SIZE_256BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
|
||||
{OPD(3, 0b01, 0x19), 1, X86InstInfo{"VEXTRACTF128", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_256BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
|
||||
{OPD(3, 0b01, 0x1D), 1, X86InstInfo{"VCVTPS2PH", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(3, 0b01, 0x1D), 1, X86InstInfo{"VCVTPS2PH", TYPE_INST, GenFlagsSizes(SIZE_64BIT, SIZE_128BIT) | FLAGS_MODRM | FLAGS_SF_MOD_DST | FLAGS_XMM_FLAGS, 1, nullptr}},
|
||||
|
||||
{OPD(3, 0b01, 0x20), 1, X86InstInfo{"VPINSRB", TYPE_INST, GenFlagsDstSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS | FLAGS_SF_SRC_GPR, 1, nullptr}},
|
||||
{OPD(3, 0b01, 0x21), 1, X86InstInfo{"VINSERTPS", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
|
||||
@@ -452,33 +452,33 @@ std::array<X86InstInfo, MAX_VEX_TABLE_SIZE> VEXTableOps = []() consteval {
|
||||
{OPD(3, 0b01, 0x4B), 1, X86InstInfo{"VBLENDVPD", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
|
||||
{OPD(3, 0b01, 0x4C), 1, X86InstInfo{"VPBLENDVB", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_VEX_1ST_SRC | FLAGS_XMM_FLAGS, 1, nullptr}},
|
||||
|
||||
{OPD(3, 0b01, 0x5C), 1, X86InstInfo{"VFMADDSUBPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(3, 0b01, 0x5D), 1, X86InstInfo{"VFMADDSUBPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(3, 0b01, 0x5E), 1, X86InstInfo{"VFMSUBADDPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(3, 0b01, 0x5F), 1, X86InstInfo{"VFMSUBADDPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(3, 0b01, 0x5C), 1, X86InstInfo{"VFMADDSUBPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, ///< FMA4
|
||||
{OPD(3, 0b01, 0x5D), 1, X86InstInfo{"VFMADDSUBPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, ///< FMA4
|
||||
{OPD(3, 0b01, 0x5E), 1, X86InstInfo{"VFMSUBADDPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, ///< FMA4
|
||||
{OPD(3, 0b01, 0x5F), 1, X86InstInfo{"VFMSUBADDPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, ///< FMA4
|
||||
|
||||
{OPD(3, 0b01, 0x60), 1, X86InstInfo{"VPCMPESTRM", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
|
||||
{OPD(3, 0b01, 0x61), 1, X86InstInfo{"VPCMPESTRI", TYPE_INST, GenFlagsSizes(SIZE_128BIT, SIZE_32BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
|
||||
{OPD(3, 0b01, 0x62), 1, X86InstInfo{"VPCMPISTRM", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
|
||||
{OPD(3, 0b01, 0x63), 1, X86InstInfo{"VPCMPISTRI", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
|
||||
|
||||
{OPD(3, 0b01, 0x68), 1, X86InstInfo{"VFMADDPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(3, 0b01, 0x69), 1, X86InstInfo{"VFMADDPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(3, 0b01, 0x6A), 1, X86InstInfo{"VFMADDSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(3, 0b01, 0x6B), 1, X86InstInfo{"VFMADDSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(3, 0b01, 0x6C), 1, X86InstInfo{"VFMSUBPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(3, 0b01, 0x6D), 1, X86InstInfo{"VFMSUBPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(3, 0b01, 0x6E), 1, X86InstInfo{"VFMSUBSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(3, 0b01, 0x6F), 1, X86InstInfo{"VFMSUBSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(3, 0b01, 0x68), 1, X86InstInfo{"VFMADDPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, ///< FMA4
|
||||
{OPD(3, 0b01, 0x69), 1, X86InstInfo{"VFMADDPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, ///< FMA4
|
||||
{OPD(3, 0b01, 0x6A), 1, X86InstInfo{"VFMADDSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, ///< FMA4
|
||||
{OPD(3, 0b01, 0x6B), 1, X86InstInfo{"VFMADDSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, ///< FMA4
|
||||
{OPD(3, 0b01, 0x6C), 1, X86InstInfo{"VFMSUBPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, ///< FMA4
|
||||
{OPD(3, 0b01, 0x6D), 1, X86InstInfo{"VFMSUBPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, ///< FMA4
|
||||
{OPD(3, 0b01, 0x6E), 1, X86InstInfo{"VFMSUBSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, ///< FMA4
|
||||
{OPD(3, 0b01, 0x6F), 1, X86InstInfo{"VFMSUBSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, ///< FMA4
|
||||
|
||||
{OPD(3, 0b01, 0x78), 1, X86InstInfo{"VFNMADDPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(3, 0b01, 0x79), 1, X86InstInfo{"VFNMADDPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(3, 0b01, 0x7A), 1, X86InstInfo{"VFNMADDSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(3, 0b01, 0x7B), 1, X86InstInfo{"VFNMADDSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(3, 0b01, 0x7C), 1, X86InstInfo{"VFNMSUBPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(3, 0b01, 0x7D), 1, X86InstInfo{"VFNMSUBPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(3, 0b01, 0x7E), 1, X86InstInfo{"VFNMSUBSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(3, 0b01, 0x7F), 1, X86InstInfo{"VFNMSUBSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}},
|
||||
{OPD(3, 0b01, 0x78), 1, X86InstInfo{"VFNMADDPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, ///< FMA4
|
||||
{OPD(3, 0b01, 0x79), 1, X86InstInfo{"VFNMADDPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, ///< FMA4
|
||||
{OPD(3, 0b01, 0x7A), 1, X86InstInfo{"VFNMADDSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, ///< FMA4
|
||||
{OPD(3, 0b01, 0x7B), 1, X86InstInfo{"VFNMADDSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, ///< FMA4
|
||||
{OPD(3, 0b01, 0x7C), 1, X86InstInfo{"VFNMSUBPS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, ///< FMA4
|
||||
{OPD(3, 0b01, 0x7D), 1, X86InstInfo{"VFNMSUBPD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, ///< FMA4
|
||||
{OPD(3, 0b01, 0x7E), 1, X86InstInfo{"VFNMSUBSS", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, ///< FMA4
|
||||
{OPD(3, 0b01, 0x7F), 1, X86InstInfo{"VFNMSUBSD", TYPE_UNDEC, FLAGS_NONE, 0, nullptr}}, ///< FMA4
|
||||
|
||||
{OPD(3, 0b01, 0xDF), 1, X86InstInfo{"VAESKEYGENASSIST", TYPE_INST, GenFlagsSameSize(SIZE_128BIT) | FLAGS_MODRM | FLAGS_XMM_FLAGS, 1, nullptr}},
|
||||
|
||||
|
||||
@@ -27,7 +27,7 @@ constexpr uint32_t FLAG_LOCK = (1 << 2);
|
||||
constexpr uint32_t FLAG_LEGACY_PREFIX = (1 << 3);
|
||||
constexpr uint32_t FLAG_REX_PREFIX = (1 << 4);
|
||||
constexpr uint32_t FLAG_VSIB_BYTE = (1 << 5);
|
||||
// Hole where 1 << 6 is
|
||||
constexpr uint32_t FLAG_OPTION_AVX_W = (1 << 6);
|
||||
constexpr uint32_t FLAG_REX_WIDENING = (1 << 7);
|
||||
constexpr uint32_t FLAG_REX_XGPR_B = (1 << 8);
|
||||
constexpr uint32_t FLAG_REX_XGPR_X = (1 << 9);
|
||||
@@ -137,6 +137,10 @@ struct DecodedOperand {
|
||||
bool IsSIB() const {
|
||||
return Type == OpType::SIB;
|
||||
}
|
||||
uint64_t Literal() const {
|
||||
LOGMAN_THROW_A_FMT(IsLiteral(), "Precondition: must be a literal");
|
||||
return Data.Literal.Value;
|
||||
}
|
||||
|
||||
union TypeUnion {
|
||||
struct GPRType {
|
||||
|
||||
@@ -227,8 +227,7 @@ struct ThunkHandler_impl final : public ThunkHandler {
|
||||
const uint8_t GPRSize = CTX->GetGPRSize();
|
||||
|
||||
if (GPRSize == 8) {
|
||||
emit->_StoreRegister(emit->_Constant(Entrypoint), false, offsetof(Core::CPUState, gregs[X86State::REG_R11]), IR::GPRClass,
|
||||
IR::GPRFixedClass, GPRSize);
|
||||
emit->_StoreRegister(emit->_Constant(Entrypoint), X86State::REG_R11, IR::GPRClass, GPRSize);
|
||||
} else {
|
||||
emit->_StoreContext(GPRSize, IR::FPRClass, emit->_VCastFromGPR(8, 8, emit->_Constant(Entrypoint)), offsetof(Core::CPUState, mm[0][0]));
|
||||
}
|
||||
|
||||
@@ -59,20 +59,20 @@ IR::IRListView* AOTIRInlineEntry::GetIRData() {
|
||||
}
|
||||
|
||||
void AOTIRCaptureCacheEntry::AppendAOTIRCaptureCache(uint64_t GuestRIP, uint64_t Start, uint64_t Length, uint64_t Hash,
|
||||
FEXCore::IR::IRListView* IRList, FEXCore::IR::RegisterAllocationData* RAData) {
|
||||
const FEXCore::IR::IRListView& IRList, const FEXCore::IR::RegisterAllocationData* RAData) {
|
||||
auto Inserted = Index.emplace(GuestRIP, Stream->Offset());
|
||||
|
||||
if (Inserted.second) {
|
||||
// GuestHash
|
||||
Stream->Write((const char*)&Hash, sizeof(Hash));
|
||||
|
||||
// GuestLength
|
||||
Stream->Write((const char*)&Length, sizeof(Length));
|
||||
AOTIRInlineEntry entry {
|
||||
.GuestHash = Hash,
|
||||
.GuestLength = Length,
|
||||
};
|
||||
Stream->Write((const char*)&entry, sizeof(entry));
|
||||
|
||||
RAData->Serialize(*Stream);
|
||||
|
||||
// IRData (inline)
|
||||
IRList->Serialize(*Stream);
|
||||
IRList.Serialize(*Stream);
|
||||
}
|
||||
}
|
||||
|
||||
@@ -170,25 +170,23 @@ void AOTIRCaptureCache::FinalizeAOTIRCache() {
|
||||
stream->Write(&Zero, 1);
|
||||
}
|
||||
|
||||
// AOTIRInlineIndex
|
||||
const auto FnCount = Entry.Index.size();
|
||||
const size_t DataBase = -stream->Offset();
|
||||
|
||||
stream->Write((const char*)&FnCount, sizeof(FnCount));
|
||||
stream->Write((const char*)&DataBase, sizeof(DataBase));
|
||||
AOTIRInlineIndex index {
|
||||
.Count = Entry.Index.size(),
|
||||
.DataBase = -stream->Offset(),
|
||||
};
|
||||
stream->Write((const char*)&index, sizeof(index));
|
||||
|
||||
for (const auto& [GuestStart, DataOffset] : Entry.Index) {
|
||||
// AOTIRInlineIndexEntry
|
||||
AOTIRInlineIndexEntry entry {
|
||||
.GuestStart = GuestStart,
|
||||
.DataOffset = DataOffset,
|
||||
};
|
||||
|
||||
// GuestStart
|
||||
stream->Write((const char*)&GuestStart, sizeof(GuestStart));
|
||||
|
||||
// DataOffset
|
||||
stream->Write((const char*)&DataOffset, sizeof(DataOffset));
|
||||
stream->Write((const char*)&entry, sizeof(entry));
|
||||
}
|
||||
|
||||
// End of file header
|
||||
const auto IndexSize = FnCount * sizeof(FEXCore::IR::AOTIRInlineIndexEntry) + sizeof(DataBase) + sizeof(FnCount);
|
||||
const auto IndexSize = sizeof(AOTIRInlineIndex) + index.Count * sizeof(FEXCore::IR::AOTIRInlineIndexEntry);
|
||||
stream->Write((const char*)&IndexSize, sizeof(IndexSize));
|
||||
stream->Write(String.c_str(), ModSize);
|
||||
stream->Write((const char*)&ModSize, sizeof(ModSize));
|
||||
@@ -261,8 +259,23 @@ void AOTIRCaptureCache::WriteFilesWithCode(const Context::AOTIRCodeFileWriterFn&
|
||||
}
|
||||
}
|
||||
|
||||
AOTIRCaptureCache::PreGenerateIRFetchResult
|
||||
AOTIRCaptureCache::PreGenerateIRFetch(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, FEXCore::IR::IRListView* IRList) {
|
||||
// IRStorageBase with memory owned by IR cache
|
||||
class IRInlineStorage : public IRStorageBase {
|
||||
AOTIRInlineEntry& entry;
|
||||
public:
|
||||
IRInlineStorage(AOTIRInlineEntry& entry)
|
||||
: entry(entry) {}
|
||||
const RegisterAllocationData* RAData() override {
|
||||
return entry.GetRAData();
|
||||
}
|
||||
IRListView GetIRView() override {
|
||||
return entry.GetIRData();
|
||||
}
|
||||
};
|
||||
|
||||
|
||||
std::optional<AOTIRCaptureCache::PreGenerateIRFetchResult>
|
||||
AOTIRCaptureCache::PreGenerateIRFetch(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP) {
|
||||
auto AOTIRCacheEntry = CTX->SyscallHandler->LookupAOTIRCacheEntry(Thread, GuestRIP);
|
||||
|
||||
PreGenerateIRFetchResult Result {};
|
||||
@@ -270,7 +283,7 @@ AOTIRCaptureCache::PreGenerateIRFetch(FEXCore::Core::InternalThreadState* Thread
|
||||
if (AOTIRCacheEntry.Entry) {
|
||||
AOTIRCacheEntry.Entry->ContainsCode = true;
|
||||
|
||||
if (IRList == nullptr && CTX->Config.AOTIRLoad()) {
|
||||
if (CTX->Config.AOTIRLoad()) {
|
||||
auto Mod = AOTIRCacheEntry.Entry->Array;
|
||||
|
||||
if (Mod != nullptr) {
|
||||
@@ -281,14 +294,12 @@ AOTIRCaptureCache::PreGenerateIRFetch(FEXCore::Core::InternalThreadState* Thread
|
||||
auto MappedStart = GuestRIP;
|
||||
auto hash = XXH3_64bits((void*)MappedStart, AOTEntry->GuestLength);
|
||||
if (hash == AOTEntry->GuestHash) {
|
||||
Result.IRList = AOTEntry->GetIRData();
|
||||
Result.IR = fextl::make_unique<IRInlineStorage>(*AOTEntry);
|
||||
// LogMan::Msg::DFmt("using {} + {:x} -> {:x}\n", file->second.fileid, AOTEntry->first, GuestRIP);
|
||||
|
||||
Result.RAData = AOTEntry->GetRAData()->CreateCopy();
|
||||
Result.DebugData = new FEXCore::Core::DebugData();
|
||||
Result.StartAddr = MappedStart;
|
||||
Result.Length = AOTEntry->GuestLength;
|
||||
Result.GeneratedIR = true;
|
||||
return Result;
|
||||
} else {
|
||||
LogMan::Msg::IFmt("AOTIR: hash check failed {:x}\n", MappedStart);
|
||||
}
|
||||
@@ -299,12 +310,12 @@ AOTIRCaptureCache::PreGenerateIRFetch(FEXCore::Core::InternalThreadState* Thread
|
||||
}
|
||||
}
|
||||
|
||||
return Result;
|
||||
return std::nullopt;
|
||||
}
|
||||
|
||||
bool AOTIRCaptureCache::PostCompileCode(FEXCore::Core::InternalThreadState* Thread, void* CodePtr, uint64_t GuestRIP, uint64_t StartAddr,
|
||||
uint64_t Length, FEXCore::IR::RegisterAllocationData::UniquePtr RAData,
|
||||
FEXCore::IR::IRListView* IRList, FEXCore::Core::DebugData* DebugData, bool GeneratedIR) {
|
||||
uint64_t Length, fextl::unique_ptr<FEXCore::IR::IRStorageBase> IR,
|
||||
FEXCore::Core::DebugData* DebugData, bool GeneratedIR) {
|
||||
|
||||
// Both generated ir and LibraryJITName need a named region lookup
|
||||
if (GeneratedIR || CTX->Config.LibraryJITNaming() || CTX->Config.GDBSymbols()) {
|
||||
@@ -321,24 +332,20 @@ bool AOTIRCaptureCache::PostCompileCode(FEXCore::Core::InternalThreadState* Thre
|
||||
}
|
||||
|
||||
// Add to AOT cache if aot generation is enabled
|
||||
if (GeneratedIR && RAData && (CTX->Config.AOTIRCapture() || CTX->Config.AOTIRGenerate())) {
|
||||
if (GeneratedIR && IR->RAData() && (CTX->Config.AOTIRCapture() || CTX->Config.AOTIRGenerate())) {
|
||||
|
||||
auto hash = XXH3_64bits((void*)StartAddr, Length);
|
||||
|
||||
auto LocalRIP = GuestRIP - AOTIRCacheEntry.VAFileStart;
|
||||
auto LocalStartAddr = StartAddr - AOTIRCacheEntry.VAFileStart;
|
||||
auto FileId = AOTIRCacheEntry.Entry->FileId;
|
||||
// The underlying pointer and the unique_ptr deleter for RAData must
|
||||
// be marshalled separately to the lambda below. Otherwise, the
|
||||
// lambda can't be used as an std::function due to being non-copyable
|
||||
auto RADataCopy = RAData->CreateCopy();
|
||||
auto RADataCopyDeleter = RADataCopy.get_deleter();
|
||||
auto IRListCopy = IRList->CreateCopy();
|
||||
|
||||
// The lambda is converted to std::function. This is tricky to refactor so it doesn't allocate memory through glibc.
|
||||
// NOTE: unique_ptr must be passed as a raw pointer since std::function requires lambda captures to be copyable
|
||||
FEXCore::Allocator::YesIKnowImNotSupposedToUseTheGlibcAllocator glibc;
|
||||
AOTIRCaptureCacheWriteoutQueue_Append(
|
||||
[this, LocalRIP, LocalStartAddr, Length, hash, IRListCopy, RADataCopy = RADataCopy.release(), RADataCopyDeleter, FileId]() {
|
||||
AOTIRCaptureCacheWriteoutQueue_Append([this, LocalRIP, LocalStartAddr, Length, hash, IRRaw = IR.release(), FileId]() {
|
||||
fextl::unique_ptr<FEXCore::IR::IRStorageBase> IR(IRRaw);
|
||||
|
||||
// It is guaranteed via AOTIRCaptureCacheWriteoutLock and AOTIRCaptureCacheWriteoutFlusing that this will not run concurrently
|
||||
// Memory coherency is guaranteed via AOTIRCaptureCacheWriteoutLock
|
||||
|
||||
@@ -349,9 +356,7 @@ bool AOTIRCaptureCache::PostCompileCode(FEXCore::Core::InternalThreadState* Thre
|
||||
uint64_t tag = FEXCore::IR::AOTIR_COOKIE;
|
||||
AotFile->Stream->Write(&tag, sizeof(tag));
|
||||
}
|
||||
AotFile->AppendAOTIRCaptureCache(LocalRIP, LocalStartAddr, Length, hash, IRListCopy, RADataCopy);
|
||||
RADataCopyDeleter(RADataCopy);
|
||||
delete IRListCopy;
|
||||
AotFile->AppendAOTIRCaptureCache(LocalRIP, LocalStartAddr, Length, hash, IR->GetIRView(), IR->RAData());
|
||||
});
|
||||
|
||||
if (CTX->Config.AOTIRGenerate()) {
|
||||
@@ -366,9 +371,6 @@ bool AOTIRCaptureCache::PostCompileCode(FEXCore::Core::InternalThreadState* Thre
|
||||
if (GeneratedIR) {
|
||||
// If the IR doesn't need to be retained then we can just delete it now
|
||||
delete DebugData;
|
||||
if (IRList->IsCopy()) {
|
||||
delete IRList;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -2,6 +2,7 @@
|
||||
#pragma once
|
||||
|
||||
#include "Interface/IR/RegisterAllocationData.h"
|
||||
#include "Interface/IR/IntrusiveIRList.h"
|
||||
|
||||
#include <FEXCore/Config/Config.h>
|
||||
#include <FEXCore/fextl/map.h>
|
||||
@@ -74,8 +75,8 @@ struct AOTIRCaptureCacheEntry {
|
||||
fextl::unique_ptr<FEXCore::Context::AOTIRWriter> Stream;
|
||||
fextl::map<uint64_t, uint64_t> Index;
|
||||
|
||||
void AppendAOTIRCaptureCache(uint64_t GuestRIP, uint64_t Start, uint64_t Length, uint64_t Hash, FEXCore::IR::IRListView* IRList,
|
||||
FEXCore::IR::RegisterAllocationData* RAData);
|
||||
void AppendAOTIRCaptureCache(uint64_t GuestRIP, uint64_t Start, uint64_t Length, uint64_t Hash, const FEXCore::IR::IRListView& IRList,
|
||||
const FEXCore::IR::RegisterAllocationData* RAData);
|
||||
};
|
||||
|
||||
struct AOTIRCacheEntry {
|
||||
@@ -103,19 +104,16 @@ public:
|
||||
void WriteFilesWithCode(const Context::AOTIRCodeFileWriterFn& Writer);
|
||||
|
||||
struct PreGenerateIRFetchResult {
|
||||
FEXCore::IR::IRListView* IRList {};
|
||||
FEXCore::IR::RegisterAllocationData::UniquePtr RAData {};
|
||||
fextl::unique_ptr<IRStorageBase> IR;
|
||||
FEXCore::Core::DebugData* DebugData {};
|
||||
uint64_t StartAddr {};
|
||||
uint64_t Length {};
|
||||
bool GeneratedIR {};
|
||||
};
|
||||
[[nodiscard]]
|
||||
PreGenerateIRFetchResult PreGenerateIRFetch(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP, FEXCore::IR::IRListView* IRList);
|
||||
std::optional<PreGenerateIRFetchResult> PreGenerateIRFetch(FEXCore::Core::InternalThreadState* Thread, uint64_t GuestRIP);
|
||||
|
||||
bool PostCompileCode(FEXCore::Core::InternalThreadState* Thread, void* CodePtr, uint64_t GuestRIP, uint64_t StartAddr, uint64_t Length,
|
||||
FEXCore::IR::RegisterAllocationData::UniquePtr RAData, FEXCore::IR::IRListView* IRList,
|
||||
FEXCore::Core::DebugData* DebugData, bool GeneratedIR);
|
||||
fextl::unique_ptr<FEXCore::IR::IRStorageBase> IR, FEXCore::Core::DebugData* DebugData, bool GeneratedIR);
|
||||
|
||||
AOTIRCacheEntry* LoadAOTIRCacheEntry(const fextl::string& filename);
|
||||
void UnloadAOTIRCacheEntry(AOTIRCacheEntry* Entry);
|
||||
|
||||
@@ -360,6 +360,17 @@ static_assert(std::is_trivially_copyable_v<OrderedNode>);
|
||||
static_assert(offsetof(OrderedNode, Header) == 0);
|
||||
static_assert(sizeof(OrderedNode) == (sizeof(OrderedNodeHeader) + sizeof(uint32_t)));
|
||||
|
||||
// This is temporary. We are transitioning away from OrderedNode's in favour of
|
||||
// flat Ref words. To ease porting, we have this typedef. Eventually OrderedNode
|
||||
// will be removed and this typedef will be replaced by something like:
|
||||
//
|
||||
// struct Ref {
|
||||
// uint Flags : 1;
|
||||
// uint ID : 23;
|
||||
// uint Reg : 8;
|
||||
// };
|
||||
using Ref = OrderedNode*;
|
||||
|
||||
struct RegisterClassType final {
|
||||
using value_type = uint32_t;
|
||||
|
||||
@@ -656,7 +667,6 @@ bool IsFragmentExit(FEXCore::IR::IROps Op);
|
||||
bool IsBlockExit(FEXCore::IR::IROps Op);
|
||||
|
||||
void Dump(fextl::stringstream* out, const IRListView* IR, IR::RegisterAllocationData* RAData);
|
||||
fextl::unique_ptr<IREmitter> Parse(FEXCore::Utils::IntrusivePooledAllocator& ThreadAllocator, fextl::stringstream& MapsStream);
|
||||
} // namespace FEXCore::IR
|
||||
|
||||
template<>
|
||||
|
||||
@@ -231,6 +231,17 @@
|
||||
],
|
||||
"HasSideEffects": true
|
||||
},
|
||||
"GPR = PushRoundingMode u8:$RoundMode": {
|
||||
"Desc": ["Override the current rounding mode options for the thread, returning old FPCR"
|
||||
],
|
||||
"DestSize": "8",
|
||||
"HasSideEffects": true
|
||||
},
|
||||
"PopRoundingMode GPR:$FPCR": {
|
||||
"Desc": ["Resets rounding mode after PushRoundingMode operation"
|
||||
],
|
||||
"HasSideEffects": true
|
||||
},
|
||||
"Print SSA:$Value": {
|
||||
"HasSideEffects": true,
|
||||
"Desc": ["Debug operation that prints an SSA value to the console",
|
||||
@@ -340,23 +351,38 @@
|
||||
"EmitValidation": [
|
||||
"Size == FEXCore::IR::OpSize::i64Bit || Size == FEXCore::IR::OpSize::i128Bit"
|
||||
]
|
||||
},
|
||||
|
||||
"GPR = Copy GPR:$Source": {
|
||||
"Desc": ["GPR copy, generated by RA to split live ranges"],
|
||||
"DestSize": "8"
|
||||
},
|
||||
|
||||
"GPR = Swap1 GPR:$A, GPR:$B": {
|
||||
"Desc": ["GPR swap part 1, generated by RA. Returns value of first source.",
|
||||
"Destination must be second GPR."],
|
||||
"DestSize": "8"
|
||||
},
|
||||
|
||||
"GPR = Swap2": {
|
||||
"Desc": ["GPR swap part 2, generated by RA. Returns source source.",
|
||||
"Must immediately succeed Swap1 with no intervening instructions",
|
||||
"Kludge to workaround single destination restriction on IR",
|
||||
"Hopefully temporary"],
|
||||
"DestSize": "8"
|
||||
}
|
||||
},
|
||||
"StaticRA": {
|
||||
"SSA = LoadRegister i1:$IsAlias, u32:$Offset, RegisterClass:$Class, RegisterClass:$StaticClass, u8:#Size": {
|
||||
"Desc": ["Loads a value from the static-ra context with offset",
|
||||
"Dest = Ctx[Offset]"
|
||||
],
|
||||
"SSA = LoadRegister u32:$Reg, RegisterClass:$Class, u8:#Size": {
|
||||
"Desc": ["Loads a value from the given register",
|
||||
"Size must match the execution mode."],
|
||||
"DestSize": "Size"
|
||||
},
|
||||
|
||||
"StoreRegister SSA:$Value, i1:$IsPrewrite, u32:$Offset, RegisterClass:$Class, RegisterClass:$StaticClass, u8:#Size": {
|
||||
"StoreRegister SSA:$Value, u32:$Reg, RegisterClass:$Class, u8:#Size": {
|
||||
"HasSideEffects": true,
|
||||
"Desc": ["Stores a value to the static-ra context with offset",
|
||||
"Ctx[Offset] = Value",
|
||||
"Zero Extends if value's type is too small",
|
||||
"Truncates if value's type is too large"
|
||||
],
|
||||
"Desc": ["Stores a value to a given register.",
|
||||
"Size must match the execution mode."],
|
||||
"DestSize": "Size",
|
||||
"EmitValidation": [
|
||||
"WalkFindRegClass($Value) == $Class"
|
||||
@@ -530,10 +556,25 @@
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / ElementSize"
|
||||
},
|
||||
"FPR = VLoadVectorGatherMasked u8:#RegisterSize, u8:#ElementSize, FPR:$Incoming, FPR:$Mask, GPR:$AddrBase, FPR:$VectorIndexLow, FPR:$VectorIndexHigh, u8:$VectorIndexElementSize, u8:$OffsetScale, u8:$DataElementOffsetStart, u8:$IndexElementOffsetStart": {
|
||||
"Desc": [
|
||||
"Does a masked load similar to VPGATHERD* where the upper bit of each element",
|
||||
"determines whether or not that element will be loaded from memory.",
|
||||
"Most of VSIB encoding is passed directly through to the IR operation."
|
||||
],
|
||||
"TiedSource": 0,
|
||||
"ImplicitFlagClobber": true,
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / ElementSize",
|
||||
"EmitValidation": [
|
||||
"$VectorIndexElementSize == OpSize::i32Bit || $VectorIndexElementSize == OpSize::i64Bit"
|
||||
]
|
||||
},
|
||||
"FPR = VLoadVectorElement u8:#RegisterSize, u8:#ElementSize, FPR:$DstSrc, u8:$Index, GPR:$Addr": {
|
||||
"Desc": ["Does a memory load to a single element of a vector.",
|
||||
"Leaves the rest of the vector's data intact.",
|
||||
"Matches arm64 ld1 semantics"],
|
||||
"TiedSource": 0,
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / ElementSize"
|
||||
},
|
||||
@@ -555,6 +596,7 @@
|
||||
"The address is decremented by the value size while.",
|
||||
"The return value size is the size of the current operating mode"
|
||||
],
|
||||
"TiedSource": 1,
|
||||
"HasSideEffects": true,
|
||||
"DestSize": "Size"
|
||||
},
|
||||
@@ -565,7 +607,7 @@
|
||||
"HasSideEffects": true,
|
||||
"DestSize": "8"
|
||||
},
|
||||
"GPRPair = MemCpy i1:$IsAtomic, u8:$Size, GPR:$PrefixDest, GPR:$PrefixSrc, GPR:$AddrDest, GPR:$AddrSrc, GPR:$Length, GPR:$Direction": {
|
||||
"GPRPair = MemCpy i1:$IsAtomic, u8:$Size, GPR:$Dest, GPR:$Src, GPR:$Length, GPR:$Direction": {
|
||||
"Desc": ["Duplicates behaviour of x86 MOVS repeat",
|
||||
"Returns the final addresses of destination and src addresses after they have been incremented or decremented"
|
||||
],
|
||||
@@ -612,6 +654,30 @@
|
||||
],
|
||||
"HasSideEffects": true,
|
||||
"DestSize": "8"
|
||||
},
|
||||
"VStoreNonTemporal u8:#RegisterSize, FPR:$Value, GPR:$Addr, i8:$Offset": {
|
||||
"Desc": ["Does a non-temporal memory store of a vector.",
|
||||
"Matches arm64 SVE stnt1b semantics.",
|
||||
"Specifically weak-memory model ordered to match x86 non-temporal stores."
|
||||
],
|
||||
"HasSideEffects": true,
|
||||
"DestSize": "RegisterSize",
|
||||
"EmitValidation": [
|
||||
"_Offset % RegisterSize == 0",
|
||||
"RegisterSize == FEXCore::IR::OpSize::i128Bit || RegisterSize == FEXCore::IR::OpSize::i256Bit"
|
||||
]
|
||||
},
|
||||
"VStoreNonTemporalPair u8:#RegisterSize, FPR:$ValueLow, FPR:$ValueHigh, GPR:$Addr, i8:$Offset": {
|
||||
"Desc": ["Does a non-temporal memory store of two vector registers.",
|
||||
"Matches arm64 stnp semantics.",
|
||||
"Specifically weak-memory model ordered to match x86 non-temporal stores."
|
||||
],
|
||||
"HasSideEffects": true,
|
||||
"DestSize": "RegisterSize",
|
||||
"EmitValidation": [
|
||||
"_Offset % RegisterSize == 0",
|
||||
"RegisterSize == FEXCore::IR::OpSize::i128Bit"
|
||||
]
|
||||
}
|
||||
},
|
||||
"Atomic": {
|
||||
@@ -986,7 +1052,8 @@
|
||||
},
|
||||
"GPR = AddShift OpSize:#Size, GPR:$Src1, GPR:$Src2, ShiftType:$Shift{ShiftType::LSL}, u8:$ShiftAmount{0}": {
|
||||
"Desc": [ "Integer Add with shifted register",
|
||||
"Will truncate to 64 or 32bits"
|
||||
"Will truncate to 64 or 32bits",
|
||||
"Dest = Src1 + (Src2 << ShiftAmount)"
|
||||
],
|
||||
"DestSize": "Size",
|
||||
"EmitValidation": [
|
||||
@@ -1180,6 +1247,7 @@
|
||||
"Desc": ["Integer binary and"
|
||||
],
|
||||
"DestSize": "Size",
|
||||
"TiedSource": 0,
|
||||
"HasSideEffects": true
|
||||
},
|
||||
"GPR = Andn OpSize:#Size, GPR:$Src1, GPR:$Src2": {
|
||||
@@ -1320,6 +1388,7 @@
|
||||
"The bitfield is copied in to Dest[(Width + lsb):lsb]"
|
||||
],
|
||||
"DestSize": "Size",
|
||||
"TiedSource": 0,
|
||||
"EmitValidation": [
|
||||
"Size == FEXCore::IR::OpSize::i32Bit || Size == FEXCore::IR::OpSize::i64Bit"
|
||||
]
|
||||
@@ -1331,6 +1400,7 @@
|
||||
"The bitfield is copied in to Dest[Width:0]"
|
||||
],
|
||||
"DestSize": "Size",
|
||||
"TiedSource": 0,
|
||||
"EmitValidation": [
|
||||
"Size == FEXCore::IR::OpSize::i32Bit || Size == FEXCore::IR::OpSize::i64Bit"
|
||||
]
|
||||
@@ -1761,29 +1831,35 @@
|
||||
"NumElements": "RegisterSize / ElementSize"
|
||||
},
|
||||
"FPR = VShlI u8:#RegisterSize, u8:#ElementSize, FPR:$Vector, u8:$BitShift": {
|
||||
"TiedSource": 0,
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / ElementSize"
|
||||
},
|
||||
"FPR = VUShrI u8:#RegisterSize, u8:#ElementSize, FPR:$Vector, u8:$BitShift": {
|
||||
"TiedSource": 0,
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / ElementSize"
|
||||
},
|
||||
"FPR = VUShraI u8:#RegisterSize, u8:#ElementSize, FPR:$DestVector, FPR:$Vector, u8:$BitShift": {
|
||||
"TiedSource": 0,
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / ElementSize"
|
||||
},
|
||||
"FPR = VSShrI u8:#RegisterSize, u8:#ElementSize, FPR:$Vector, u8:$BitShift": {
|
||||
"TiedSource": 0,
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / ElementSize"
|
||||
},
|
||||
|
||||
"FPR = VUShrNI u8:#RegisterSize, u8:#ElementSize, FPR:$Vector, u8:$BitShift": {
|
||||
"TiedSource": 0,
|
||||
"Desc": "Unsigned shifts right each element and then narrows to the next lower element size",
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / (ElementSize >> 1)"
|
||||
},
|
||||
|
||||
"FPR = VUShrNI2 u8:#RegisterSize, u8:#ElementSize, FPR:$VectorLower, FPR:$VectorUpper, u8:$BitShift": {
|
||||
"TiedSource": 0,
|
||||
"Desc": ["Unsigned shifts right each element and then narrows to the next lower element size",
|
||||
"Inserts results in to the high elements of the first argument"
|
||||
],
|
||||
@@ -1815,10 +1891,12 @@
|
||||
"NumElements": "RegisterSize / (ElementSize << 1)"
|
||||
},
|
||||
"FPR = VSQXTN u8:#RegisterSize, u8:#ElementSize, FPR:$Vector": {
|
||||
"TiedSource": 0,
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / (ElementSize >> 1)"
|
||||
},
|
||||
"FPR = VSQXTN2 u8:#RegisterSize, u8:#ElementSize, FPR:$VectorLower, FPR:$VectorUpper": {
|
||||
"TiedSource": 0,
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / (ElementSize >> 1)"
|
||||
},
|
||||
@@ -1846,6 +1924,7 @@
|
||||
"Desc": ["Signed rounding shift right by immediate",
|
||||
"Exactly matching Arm64 srshr semantics"
|
||||
],
|
||||
"TiedSource": 0,
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / ElementSize"
|
||||
},
|
||||
@@ -1853,6 +1932,7 @@
|
||||
"Desc": ["Signed satuating shift left by immediate",
|
||||
"Exactly matching Arm64 sqshl semantics"
|
||||
],
|
||||
"TiedSource": 0,
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / ElementSize"
|
||||
},
|
||||
@@ -1886,7 +1966,7 @@
|
||||
"NumElements": "RegisterSize / ElementSize"
|
||||
},
|
||||
|
||||
"FPR = VBic u8:#RegisterSize, u8:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
|
||||
"FPR = VAndn u8:#RegisterSize, u8:#ElementSize, FPR:$Vector1, FPR:$Vector2": {
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / ElementSize"
|
||||
},
|
||||
@@ -2061,42 +2141,52 @@
|
||||
"NumElements": "RegisterSize / (ElementSize << 1)"
|
||||
},
|
||||
"FPR = VUShl u8:#RegisterSize, u8:#ElementSize, FPR:$Vector, FPR:$ShiftVector, i1:$RangeCheck": {
|
||||
"TiedSource": 0,
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / ElementSize"
|
||||
},
|
||||
"FPR = VUShr u8:#RegisterSize, u8:#ElementSize, FPR:$Vector, FPR:$ShiftVector, i1:$RangeCheck": {
|
||||
"TiedSource": 0,
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / ElementSize"
|
||||
},
|
||||
"FPR = VSShr u8:#RegisterSize, u8:#ElementSize, FPR:$Vector, FPR:$ShiftVector, i1:$RangeCheck": {
|
||||
"TiedSource": 0,
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / ElementSize"
|
||||
},
|
||||
"FPR = VUShlS u8:#RegisterSize, u8:#ElementSize, FPR:$Vector, FPR:$ShiftScalar": {
|
||||
"TiedSource": 0,
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / ElementSize"
|
||||
},
|
||||
"FPR = VUShrS u8:#RegisterSize, u8:#ElementSize, FPR:$Vector, FPR:$ShiftScalar": {
|
||||
"TiedSource": 0,
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / ElementSize"
|
||||
},
|
||||
"FPR = VSShrS u8:#RegisterSize, u8:#ElementSize, FPR:$Vector, FPR:$ShiftScalar": {
|
||||
"TiedSource": 0,
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / ElementSize"
|
||||
},
|
||||
"FPR = VUShrSWide u8:#RegisterSize, u8:#ElementSize, FPR:$Vector, FPR:$ShiftScalar": {
|
||||
"TiedSource": 0,
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / ElementSize"
|
||||
},
|
||||
"FPR = VSShrSWide u8:#RegisterSize, u8:#ElementSize, FPR:$Vector, FPR:$ShiftScalar": {
|
||||
"TiedSource": 0,
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / ElementSize"
|
||||
},
|
||||
"FPR = VUShlSWide u8:#RegisterSize, u8:#ElementSize, FPR:$Vector, FPR:$ShiftScalar": {
|
||||
"TiedSource": 0,
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / ElementSize"
|
||||
},
|
||||
"FPR = VInsElement u8:#RegisterSize, u8:#ElementSize, u8:$DestIdx, u8:$SrcIdx, FPR:$DestVector, FPR:$SrcVector": {
|
||||
"TiedSource": 0,
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / ElementSize"
|
||||
},
|
||||
@@ -2183,6 +2273,7 @@
|
||||
"Table is always treated as a 128bit register",
|
||||
"Indices matches destination size. Either 64bit or 128bit"
|
||||
],
|
||||
"TiedSource": 0,
|
||||
"DestSize": "RegisterSize"
|
||||
},
|
||||
"FPR = VBSL u8:#RegisterSize, FPR:$VectorMask, FPR:$VectorTrue, FPR:$VectorFalse": {
|
||||
@@ -2220,6 +2311,42 @@
|
||||
"FPR = VFCADD u8:#RegisterSize, u8:#ElementSize, FPR:$Vector1, FPR:$Vector2, u16:$Rotate": {
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / ElementSize"
|
||||
},
|
||||
"FPR = VFMLA u8:#RegisterSize, u8:#ElementSize, FPR:$Vector1, FPR:$Vector2, FPR:$Addend": {
|
||||
"Desc": [
|
||||
"Dest = (Vector1 * Vector2) + Addend",
|
||||
"This explicitly matches x86 FMA semantics because ARM semantics are mind-bending."
|
||||
],
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / ElementSize",
|
||||
"TiedSource": 2
|
||||
},
|
||||
"FPR = VFMLS u8:#RegisterSize, u8:#ElementSize, FPR:$Vector1, FPR:$Vector2, FPR:$Addend": {
|
||||
"Desc": [
|
||||
"Dest = (Vector1 * Vector2) - Addend",
|
||||
"This explicitly matches x86 FMA semantics because ARM semantics are mind-bending."
|
||||
],
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / ElementSize",
|
||||
"TiedSource": 2
|
||||
},
|
||||
"FPR = VFNMLA u8:#RegisterSize, u8:#ElementSize, FPR:$Vector1, FPR:$Vector2, FPR:$Addend": {
|
||||
"Desc": [
|
||||
"Dest = (-Vector1 * Vector2) + Addend",
|
||||
"This explicitly matches x86 FMA semantics because ARM semantics are mind-bending."
|
||||
],
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / ElementSize",
|
||||
"TiedSource": 2
|
||||
},
|
||||
"FPR = VFNMLS u8:#RegisterSize, u8:#ElementSize, FPR:$Vector1, FPR:$Vector2, FPR:$Addend": {
|
||||
"Desc": [
|
||||
"Dest = (-Vector1 * Vector2) - Addend",
|
||||
"This explicitly matches x86 FMA semantics because ARM semantics are mind-bending."
|
||||
],
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / ElementSize",
|
||||
"TiedSource": 2
|
||||
}
|
||||
},
|
||||
"Conv": {
|
||||
@@ -2272,6 +2399,32 @@
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / DestElementSize"
|
||||
},
|
||||
|
||||
"FPR = VFCVTL2 u8:#RegisterSize, u8:#ElementSize, FPR:$Vector": {
|
||||
"Desc": [
|
||||
"Vector op: Converts float from source element size to destination size (fp32->fp64)",
|
||||
"Selecting from the high half of the register."
|
||||
],
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / (ElementSize << 1)",
|
||||
"EmitValidation": [
|
||||
"RegisterSize != FEXCore::IR::OpSize::i256Bit && \"What does 256-bit mean in this context?\""
|
||||
]
|
||||
},
|
||||
"FPR = VFCVTN2 u8:#RegisterSize, u8:#ElementSize, FPR:$VectorLower, FPR:$VectorUpper": {
|
||||
"TiedSource": 0,
|
||||
"Desc": [
|
||||
"Vector op: Converts float from source element size and inserting in to the high bits.",
|
||||
"Bottom half is untouched",
|
||||
"Narrowing to the element size below what is passed in.",
|
||||
"F64->F32, F32->F16"
|
||||
],
|
||||
"DestSize": "RegisterSize",
|
||||
"NumElements": "RegisterSize / (ElementSize >> 1)",
|
||||
"EmitValidation": [
|
||||
"RegisterSize != FEXCore::IR::OpSize::i256Bit && \"What does 256-bit mean in this context?\""
|
||||
]
|
||||
},
|
||||
"FPR = Vector_FToI u8:#RegisterSize, u8:#ElementSize, FPR:$Vector, RoundType:$Round": {
|
||||
"Desc": ["Vector op: Rounds float to integral",
|
||||
"Rounding mode determined by argument"
|
||||
|
||||
@@ -183,6 +183,14 @@ static void PrintArg(fextl::stringstream* out, [[maybe_unused]] const IRListView
|
||||
return "addsubpd_invert";
|
||||
case NamedVectorConstant::NAMED_VECTOR_PADDSUBPD_INVERT_UPPER:
|
||||
return "addsubpd_invert_upper";
|
||||
case NamedVectorConstant::NAMED_VECTOR_PSUBADDPS_INVERT:
|
||||
return "subaddps_invert";
|
||||
case NamedVectorConstant::NAMED_VECTOR_PSUBADDPS_INVERT_UPPER:
|
||||
return "subaddps_invert_upper";
|
||||
case NamedVectorConstant::NAMED_VECTOR_PSUBADDPD_INVERT:
|
||||
return "subaddpd_invert";
|
||||
case NamedVectorConstant::NAMED_VECTOR_PSUBADDPD_INVERT_UPPER:
|
||||
return "subaddpd_invert_upper";
|
||||
case NamedVectorConstant::NAMED_VECTOR_MOVMSKPS_SHIFT:
|
||||
return "movmskps_shift";
|
||||
case NamedVectorConstant::NAMED_VECTOR_AESKEYGENASSIST_SWIZZLE:
|
||||
|
||||
@@ -34,7 +34,7 @@ bool IsBlockExit(FEXCore::IR::IROps Op) {
|
||||
}
|
||||
}
|
||||
|
||||
FEXCore::IR::RegisterClassType IREmitter::WalkFindRegClass(OrderedNode* Node) {
|
||||
FEXCore::IR::RegisterClassType IREmitter::WalkFindRegClass(Ref Node) {
|
||||
auto Class = GetOpRegClass(Node);
|
||||
switch (Class) {
|
||||
case GPRClass:
|
||||
@@ -92,12 +92,12 @@ void IREmitter::ResetWorkingList() {
|
||||
CodeBlocks.clear();
|
||||
CurrentWriteCursor = nullptr;
|
||||
// This is necessary since we do "null" pointer checks
|
||||
InvalidNode = reinterpret_cast<OrderedNode*>(DualListData.ListAllocate(sizeof(OrderedNode)));
|
||||
InvalidNode = reinterpret_cast<Ref>(DualListData.ListAllocate(sizeof(OrderedNode)));
|
||||
memset(InvalidNode, 0, sizeof(OrderedNode));
|
||||
CurrentCodeBlock = nullptr;
|
||||
}
|
||||
|
||||
void IREmitter::ReplaceAllUsesWithRange(OrderedNode* Node, OrderedNode* NewNode, AllNodesIterator Begin, AllNodesIterator End) {
|
||||
void IREmitter::ReplaceAllUsesWithRange(Ref Node, Ref NewNode, AllNodesIterator Begin, AllNodesIterator End) {
|
||||
uintptr_t ListBegin = DualListData.ListBegin();
|
||||
auto NodeId = Node->Wrapped(ListBegin).ID();
|
||||
|
||||
@@ -122,19 +122,19 @@ void IREmitter::ReplaceAllUsesWithRange(OrderedNode* Node, OrderedNode* NewNode,
|
||||
}
|
||||
}
|
||||
|
||||
void IREmitter::ReplaceNodeArgument(OrderedNode* Node, uint8_t Arg, OrderedNode* NewArg) {
|
||||
void IREmitter::ReplaceNodeArgument(Ref Node, uint8_t Arg, Ref NewArg) {
|
||||
uintptr_t ListBegin = DualListData.ListBegin();
|
||||
uintptr_t DataBegin = DualListData.DataBegin();
|
||||
|
||||
FEXCore::IR::IROp_Header* IROp = Node->Op(DataBegin);
|
||||
OrderedNodeWrapper OldArgWrapper = IROp->Args[Arg];
|
||||
OrderedNode* OldArg = OldArgWrapper.GetNode(ListBegin);
|
||||
Ref OldArg = OldArgWrapper.GetNode(ListBegin);
|
||||
OldArg->RemoveUse();
|
||||
NewArg->AddUse();
|
||||
IROp->Args[Arg].NodeOffset = NewArg->Wrapped(ListBegin).NodeOffset;
|
||||
}
|
||||
|
||||
void IREmitter::RemoveArgUses(OrderedNode* Node) {
|
||||
void IREmitter::RemoveArgUses(Ref Node) {
|
||||
uintptr_t ListBegin = DualListData.ListBegin();
|
||||
uintptr_t DataBegin = DualListData.DataBegin();
|
||||
|
||||
@@ -147,13 +147,13 @@ void IREmitter::RemoveArgUses(OrderedNode* Node) {
|
||||
}
|
||||
}
|
||||
|
||||
void IREmitter::Remove(OrderedNode* Node) {
|
||||
void IREmitter::Remove(Ref Node) {
|
||||
RemoveArgUses(Node);
|
||||
|
||||
Node->Unlink(DualListData.ListBegin());
|
||||
}
|
||||
|
||||
IREmitter::IRPair<IROp_CodeBlock> IREmitter::CreateNewCodeBlockAfter(OrderedNode* insertAfter) {
|
||||
IREmitter::IRPair<IROp_CodeBlock> IREmitter::CreateNewCodeBlockAfter(Ref insertAfter) {
|
||||
auto OldCursor = GetWriteCursor();
|
||||
|
||||
auto CodeNode = CreateCodeNode();
|
||||
@@ -179,14 +179,14 @@ IREmitter::IRPair<IROp_CodeBlock> IREmitter::CreateNewCodeBlockAfter(OrderedNode
|
||||
return CodeNode;
|
||||
}
|
||||
|
||||
void IREmitter::SetCurrentCodeBlock(OrderedNode* Node) {
|
||||
void IREmitter::SetCurrentCodeBlock(Ref Node) {
|
||||
CurrentCodeBlock = Node;
|
||||
LOGMAN_THROW_A_FMT(Node->Op(DualListData.DataBegin())->Op == OP_CODEBLOCK, "Node wasn't codeblock. It was '{}'",
|
||||
IR::GetName(Node->Op(DualListData.DataBegin())->Op));
|
||||
SetWriteCursor(Node->Op(DualListData.DataBegin())->CW<IROp_CodeBlock>()->Begin.GetNode(DualListData.ListBegin()));
|
||||
}
|
||||
|
||||
void IREmitter::ReplaceWithConstant(OrderedNode* Node, uint64_t Value) {
|
||||
void IREmitter::ReplaceWithConstant(Ref Node, uint64_t Value) {
|
||||
auto Header = Node->Op(DualListData.DataBegin());
|
||||
|
||||
if (IRSizes[Header->Op] >= sizeof(IROp_Constant)) {
|
||||
|
||||
@@ -41,10 +41,7 @@ public:
|
||||
}
|
||||
|
||||
IRListView ViewIR() {
|
||||
return IRListView(&DualListData, false);
|
||||
}
|
||||
IRListView* CreateIRCopy() {
|
||||
return new IRListView(&DualListData, true);
|
||||
return IRListView(&DualListData);
|
||||
}
|
||||
void ResetWorkingList();
|
||||
|
||||
@@ -53,7 +50,7 @@ public:
|
||||
*
|
||||
* @{ */
|
||||
|
||||
FEXCore::IR::RegisterClassType WalkFindRegClass(OrderedNode* Node);
|
||||
FEXCore::IR::RegisterClassType WalkFindRegClass(Ref Node);
|
||||
|
||||
// These handlers add cost to the constructor and destructor
|
||||
// If it becomes an issue then blow them away
|
||||
@@ -73,14 +70,14 @@ public:
|
||||
IRPair<IROp_Jump> _Jump() {
|
||||
return _Jump(InvalidNode);
|
||||
}
|
||||
IRPair<IROp_CondJump> _CondJump(OrderedNode* ssa0, CondClassType cond = {COND_NEQ}) {
|
||||
IRPair<IROp_CondJump> _CondJump(Ref ssa0, CondClassType cond = {COND_NEQ}) {
|
||||
return _CondJump(ssa0, _Constant(0), InvalidNode, InvalidNode, cond, GetOpSize(ssa0));
|
||||
}
|
||||
IRPair<IROp_CondJump> _CondJump(OrderedNode* ssa0, OrderedNode* ssa1, OrderedNode* ssa2, CondClassType cond = {COND_NEQ}) {
|
||||
IRPair<IROp_CondJump> _CondJump(Ref ssa0, Ref ssa1, Ref ssa2, CondClassType cond = {COND_NEQ}) {
|
||||
return _CondJump(ssa0, _Constant(0), ssa1, ssa2, cond, GetOpSize(ssa0));
|
||||
}
|
||||
// TODO: Work to remove this implicit sized Select implementation.
|
||||
IRPair<IROp_Select> _Select(uint8_t Cond, OrderedNode* ssa0, OrderedNode* ssa1, OrderedNode* ssa2, OrderedNode* ssa3, uint8_t CompareSize = 0) {
|
||||
IRPair<IROp_Select> _Select(uint8_t Cond, Ref ssa0, Ref ssa1, Ref ssa2, Ref ssa3, uint8_t CompareSize = 0) {
|
||||
if (CompareSize == 0) {
|
||||
CompareSize = std::max<uint8_t>(4, std::max<uint8_t>(GetOpSize(ssa0), GetOpSize(ssa1)));
|
||||
}
|
||||
@@ -88,54 +85,53 @@ public:
|
||||
return _Select(IR::SizeToOpSize(std::max<uint8_t>(4, std::max<uint8_t>(GetOpSize(ssa2), GetOpSize(ssa3)))),
|
||||
IR::SizeToOpSize(CompareSize), CondClassType {Cond}, ssa0, ssa1, ssa2, ssa3);
|
||||
}
|
||||
IRPair<IROp_LoadMem> _LoadMem(FEXCore::IR::RegisterClassType Class, uint8_t Size, OrderedNode* ssa0, uint8_t Align = 1) {
|
||||
IRPair<IROp_LoadMem> _LoadMem(FEXCore::IR::RegisterClassType Class, uint8_t Size, Ref ssa0, uint8_t Align = 1) {
|
||||
return _LoadMem(Class, Size, ssa0, Invalid(), Align, MEM_OFFSET_SXTX, 1);
|
||||
}
|
||||
IRPair<IROp_LoadMemTSO> _LoadMemTSO(FEXCore::IR::RegisterClassType Class, uint8_t Size, OrderedNode* ssa0, uint8_t Align = 1) {
|
||||
IRPair<IROp_LoadMemTSO> _LoadMemTSO(FEXCore::IR::RegisterClassType Class, uint8_t Size, Ref ssa0, uint8_t Align = 1) {
|
||||
return _LoadMemTSO(Class, Size, ssa0, Invalid(), Align, MEM_OFFSET_SXTX, 1);
|
||||
}
|
||||
IRPair<IROp_StoreMem> _StoreMem(FEXCore::IR::RegisterClassType Class, uint8_t Size, OrderedNode* Addr, OrderedNode* Value, uint8_t Align = 1) {
|
||||
IRPair<IROp_StoreMem> _StoreMem(FEXCore::IR::RegisterClassType Class, uint8_t Size, Ref Addr, Ref Value, uint8_t Align = 1) {
|
||||
return _StoreMem(Class, Size, Value, Addr, Invalid(), Align, MEM_OFFSET_SXTX, 1);
|
||||
}
|
||||
IRPair<IROp_StoreMemTSO>
|
||||
_StoreMemTSO(FEXCore::IR::RegisterClassType Class, uint8_t Size, OrderedNode* Addr, OrderedNode* Value, uint8_t Align = 1) {
|
||||
IRPair<IROp_StoreMemTSO> _StoreMemTSO(FEXCore::IR::RegisterClassType Class, uint8_t Size, Ref Addr, Ref Value, uint8_t Align = 1) {
|
||||
return _StoreMemTSO(Class, Size, Value, Addr, Invalid(), Align, MEM_OFFSET_SXTX, 1);
|
||||
}
|
||||
OrderedNode* Invalid() {
|
||||
Ref Invalid() {
|
||||
return InvalidNode;
|
||||
}
|
||||
|
||||
void SetJumpTarget(IR::IROp_Jump* Op, OrderedNode* Target) {
|
||||
void SetJumpTarget(IR::IROp_Jump* Op, Ref Target) {
|
||||
LOGMAN_THROW_A_FMT(Target->Op(DualListData.DataBegin())->Op == OP_CODEBLOCK, "Tried setting Jump target to %{} {}",
|
||||
Target->Wrapped(DualListData.ListBegin()).ID(), IR::GetName(Target->Op(DualListData.DataBegin())->Op));
|
||||
|
||||
Op->Header.Args[0].NodeOffset = Target->Wrapped(DualListData.ListBegin()).NodeOffset;
|
||||
}
|
||||
void SetTrueJumpTarget(IR::IROp_CondJump* Op, OrderedNode* Target) {
|
||||
void SetTrueJumpTarget(IR::IROp_CondJump* Op, Ref Target) {
|
||||
LOGMAN_THROW_A_FMT(Target->Op(DualListData.DataBegin())->Op == OP_CODEBLOCK, "Tried setting CondJump target to %{} {}",
|
||||
Target->Wrapped(DualListData.ListBegin()).ID(), IR::GetName(Target->Op(DualListData.DataBegin())->Op));
|
||||
|
||||
Op->TrueBlock.NodeOffset = Target->Wrapped(DualListData.ListBegin()).NodeOffset;
|
||||
}
|
||||
void SetFalseJumpTarget(IR::IROp_CondJump* Op, OrderedNode* Target) {
|
||||
void SetFalseJumpTarget(IR::IROp_CondJump* Op, Ref Target) {
|
||||
LOGMAN_THROW_A_FMT(Target->Op(DualListData.DataBegin())->Op == OP_CODEBLOCK, "Tried setting CondJump target to %{} {}",
|
||||
Target->Wrapped(DualListData.ListBegin()).ID(), IR::GetName(Target->Op(DualListData.DataBegin())->Op));
|
||||
|
||||
Op->FalseBlock.NodeOffset = Target->Wrapped(DualListData.ListBegin()).NodeOffset;
|
||||
}
|
||||
|
||||
void SetJumpTarget(IRPair<IROp_Jump> Op, OrderedNode* Target) {
|
||||
void SetJumpTarget(IRPair<IROp_Jump> Op, Ref Target) {
|
||||
LOGMAN_THROW_A_FMT(Target->Op(DualListData.DataBegin())->Op == OP_CODEBLOCK, "Tried setting Jump target to %{} {}",
|
||||
Target->Wrapped(DualListData.ListBegin()).ID(), IR::GetName(Target->Op(DualListData.DataBegin())->Op));
|
||||
|
||||
Op.first->Header.Args[0].NodeOffset = Target->Wrapped(DualListData.ListBegin()).NodeOffset;
|
||||
}
|
||||
void SetTrueJumpTarget(IRPair<IROp_CondJump> Op, OrderedNode* Target) {
|
||||
void SetTrueJumpTarget(IRPair<IROp_CondJump> Op, Ref Target) {
|
||||
LOGMAN_THROW_A_FMT(Target->Op(DualListData.DataBegin())->Op == OP_CODEBLOCK, "Tried setting CondJump target to %{} {}",
|
||||
Target->Wrapped(DualListData.ListBegin()).ID(), IR::GetName(Target->Op(DualListData.DataBegin())->Op));
|
||||
Op.first->TrueBlock.NodeOffset = Target->Wrapped(DualListData.ListBegin()).NodeOffset;
|
||||
}
|
||||
void SetFalseJumpTarget(IRPair<IROp_CondJump> Op, OrderedNode* Target) {
|
||||
void SetFalseJumpTarget(IRPair<IROp_CondJump> Op, Ref Target) {
|
||||
LOGMAN_THROW_A_FMT(Target->Op(DualListData.DataBegin())->Op == OP_CODEBLOCK, "Tried setting CondJump target to %{} {}",
|
||||
Target->Wrapped(DualListData.ListBegin()).ID(), IR::GetName(Target->Op(DualListData.DataBegin())->Op));
|
||||
Op.first->FalseBlock.NodeOffset = Target->Wrapped(DualListData.ListBegin()).NodeOffset;
|
||||
@@ -143,12 +139,12 @@ public:
|
||||
|
||||
/** @} */
|
||||
FEXCore::IR::RegisterClassType WalkFindRegClass(OrderedNodeWrapper ssa) {
|
||||
OrderedNode* RealNode = ssa.GetNode(DualListData.ListBegin());
|
||||
Ref RealNode = ssa.GetNode(DualListData.ListBegin());
|
||||
return WalkFindRegClass(RealNode);
|
||||
}
|
||||
|
||||
bool IsValueConstant(OrderedNodeWrapper ssa, uint64_t* Constant = nullptr) {
|
||||
OrderedNode* RealNode = ssa.GetNode(DualListData.ListBegin());
|
||||
Ref RealNode = ssa.GetNode(DualListData.ListBegin());
|
||||
FEXCore::IR::IROp_Header* IROp = RealNode->Op(DualListData.DataBegin());
|
||||
if (IROp->Op == OP_CONSTANT) {
|
||||
auto Op = IROp->C<IR::IROp_Constant>();
|
||||
@@ -161,7 +157,7 @@ public:
|
||||
}
|
||||
|
||||
bool IsValueInlineConstant(OrderedNodeWrapper ssa) {
|
||||
OrderedNode* RealNode = ssa.GetNode(DualListData.ListBegin());
|
||||
Ref RealNode = ssa.GetNode(DualListData.ListBegin());
|
||||
FEXCore::IR::IROp_Header* IROp = RealNode->Op(DualListData.DataBegin());
|
||||
if (IROp->Op == OP_INLINECONSTANT) {
|
||||
return true;
|
||||
@@ -170,15 +166,15 @@ public:
|
||||
}
|
||||
|
||||
FEXCore::IR::IROp_Header* GetOpHeader(OrderedNodeWrapper ssa) {
|
||||
OrderedNode* RealNode = ssa.GetNode(DualListData.ListBegin());
|
||||
Ref RealNode = ssa.GetNode(DualListData.ListBegin());
|
||||
return RealNode->Op(DualListData.DataBegin());
|
||||
}
|
||||
|
||||
OrderedNode* UnwrapNode(OrderedNodeWrapper ssa) {
|
||||
Ref UnwrapNode(OrderedNodeWrapper ssa) {
|
||||
return ssa.GetNode(DualListData.ListBegin());
|
||||
}
|
||||
|
||||
OrderedNodeWrapper WrapNode(OrderedNode* node) {
|
||||
OrderedNodeWrapper WrapNode(Ref node) {
|
||||
return node->Wrapped(DualListData.ListBegin());
|
||||
}
|
||||
|
||||
@@ -189,23 +185,23 @@ public:
|
||||
// Overwrite a node with a constant
|
||||
// Depending on what node has been overwritten, there might be some unallocated space around the node
|
||||
// Because we are overwriting the node, we don't have to worry about update all the arguments which use it
|
||||
void ReplaceWithConstant(OrderedNode* Node, uint64_t Value);
|
||||
void ReplaceWithConstant(Ref Node, uint64_t Value);
|
||||
|
||||
void ReplaceAllUsesWithRange(OrderedNode* Node, OrderedNode* NewNode, AllNodesIterator Begin, AllNodesIterator End);
|
||||
void ReplaceAllUsesWithRange(Ref Node, Ref NewNode, AllNodesIterator Begin, AllNodesIterator End);
|
||||
|
||||
void ReplaceUsesWithAfter(OrderedNode* Node, OrderedNode* NewNode, AllNodesIterator After) {
|
||||
void ReplaceUsesWithAfter(Ref Node, Ref NewNode, AllNodesIterator After) {
|
||||
++After;
|
||||
ReplaceAllUsesWithRange(Node, NewNode, After, AllNodesIterator(DualListData.ListBegin(), DualListData.DataBegin()));
|
||||
}
|
||||
|
||||
void ReplaceUsesWithAfter(OrderedNode* Node, OrderedNode* NewNode, OrderedNode* After) {
|
||||
void ReplaceUsesWithAfter(Ref Node, Ref NewNode, Ref After) {
|
||||
auto Wrapped = After->Wrapped(DualListData.ListBegin());
|
||||
AllNodesIterator It = AllNodesIterator(DualListData.ListBegin(), DualListData.DataBegin(), Wrapped);
|
||||
|
||||
ReplaceUsesWithAfter(Node, NewNode, It);
|
||||
}
|
||||
|
||||
void ReplaceAllUsesWith(OrderedNode* Node, OrderedNode* NewNode) {
|
||||
void ReplaceAllUsesWith(Ref Node, Ref NewNode) {
|
||||
auto Start = AllNodesIterator(DualListData.ListBegin(), DualListData.DataBegin(), Node->Wrapped(DualListData.ListBegin()));
|
||||
|
||||
ReplaceAllUsesWithRange(Node, NewNode, Start, AllNodesIterator(DualListData.ListBegin(), DualListData.DataBegin()));
|
||||
@@ -220,12 +216,12 @@ public:
|
||||
}
|
||||
}
|
||||
|
||||
void ReplaceNodeArgument(OrderedNode* Node, uint8_t Arg, OrderedNode* NewArg);
|
||||
void ReplaceNodeArgument(Ref Node, uint8_t Arg, Ref NewArg);
|
||||
|
||||
void Remove(OrderedNode* Node);
|
||||
void Remove(Ref Node);
|
||||
|
||||
void SetPackedRFLAG(bool Lower8, OrderedNode* Src);
|
||||
OrderedNode* GetPackedRFLAG(bool Lower8);
|
||||
void SetPackedRFLAG(bool Lower8, Ref Src);
|
||||
Ref GetPackedRFLAG(bool Lower8);
|
||||
|
||||
void CopyData(const IREmitter& rhs) {
|
||||
LOGMAN_THROW_A_FMT(rhs.DualListData.DataBackingSize() <= DualListData.DataBackingSize(), "Trying to take ownership of data that is too "
|
||||
@@ -241,12 +237,12 @@ public:
|
||||
}
|
||||
}
|
||||
|
||||
void SetWriteCursor(OrderedNode* Node) {
|
||||
void SetWriteCursor(Ref Node) {
|
||||
CurrentWriteCursor = Node;
|
||||
}
|
||||
|
||||
// Set cursor to write before Node
|
||||
void SetWriteCursorBefore(OrderedNode* Node) {
|
||||
void SetWriteCursorBefore(Ref Node) {
|
||||
auto IR = ViewIR();
|
||||
auto Before = IR.at(Node);
|
||||
--Before;
|
||||
@@ -254,11 +250,11 @@ public:
|
||||
SetWriteCursor(std::get<0>(*Before));
|
||||
}
|
||||
|
||||
OrderedNode* GetWriteCursor() {
|
||||
Ref GetWriteCursor() {
|
||||
return CurrentWriteCursor;
|
||||
}
|
||||
|
||||
OrderedNode* GetCurrentBlock() {
|
||||
Ref GetCurrentBlock() {
|
||||
return CurrentCodeBlock;
|
||||
}
|
||||
|
||||
@@ -300,7 +296,7 @@ public:
|
||||
*
|
||||
* @{ */
|
||||
/** @} */
|
||||
void LinkCodeBlocks(OrderedNode* CodeNode, OrderedNode* Next) {
|
||||
void LinkCodeBlocks(Ref CodeNode, Ref Next) {
|
||||
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
|
||||
FEXCore::IR::IROp_CodeBlock* CurrentIROp =
|
||||
#endif
|
||||
@@ -314,17 +310,17 @@ public:
|
||||
IRPair<IROp_CodeBlock> CreateNewCodeBlockAtEnd() {
|
||||
return CreateNewCodeBlockAfter(nullptr);
|
||||
}
|
||||
IRPair<IROp_CodeBlock> CreateNewCodeBlockAfter(OrderedNode* insertAfter);
|
||||
void SetCurrentCodeBlock(OrderedNode* Node);
|
||||
IRPair<IROp_CodeBlock> CreateNewCodeBlockAfter(Ref insertAfter);
|
||||
void SetCurrentCodeBlock(Ref Node);
|
||||
|
||||
protected:
|
||||
void RemoveArgUses(OrderedNode* Node);
|
||||
void RemoveArgUses(Ref Node);
|
||||
|
||||
OrderedNode* CreateNode(IROp_Header* Op) {
|
||||
Ref CreateNode(IROp_Header* Op) {
|
||||
uintptr_t ListBegin = DualListData.ListBegin();
|
||||
size_t Size = sizeof(OrderedNode);
|
||||
void* Ptr = DualListData.ListAllocate(Size);
|
||||
OrderedNode* Node = new (Ptr) OrderedNode();
|
||||
Ref Node = new (Ptr) OrderedNode();
|
||||
Node->Header.Value.SetOffset(DualListData.DataBegin(), reinterpret_cast<uintptr_t>(Op));
|
||||
|
||||
if (CurrentWriteCursor) {
|
||||
@@ -334,15 +330,15 @@ protected:
|
||||
return Node;
|
||||
}
|
||||
|
||||
OrderedNode* GetNode(uint32_t SSANode) {
|
||||
Ref GetNode(uint32_t SSANode) {
|
||||
uintptr_t ListBegin = DualListData.ListBegin();
|
||||
OrderedNode* Node = reinterpret_cast<OrderedNode*>(ListBegin + SSANode * sizeof(OrderedNode));
|
||||
Ref Node = reinterpret_cast<Ref>(ListBegin + SSANode * sizeof(OrderedNode));
|
||||
return Node;
|
||||
}
|
||||
|
||||
OrderedNode* EmplaceOrphanedNode(OrderedNode* OldNode) {
|
||||
Ref EmplaceOrphanedNode(Ref OldNode) {
|
||||
size_t Size = sizeof(OrderedNode);
|
||||
OrderedNode* Ptr = reinterpret_cast<OrderedNode*>(DualListData.ListAllocate(Size));
|
||||
Ref Ptr = reinterpret_cast<Ref>(DualListData.ListAllocate(Size));
|
||||
memcpy(Ptr, OldNode, Size);
|
||||
return Ptr;
|
||||
}
|
||||
@@ -351,14 +347,14 @@ protected:
|
||||
// Overriden by dispatcher, stubbed for IR tests
|
||||
}
|
||||
|
||||
OrderedNode* CurrentWriteCursor = nullptr;
|
||||
Ref CurrentWriteCursor = nullptr;
|
||||
|
||||
// These could be combined with a little bit of work to be more efficient with memory usage. Isn't a big deal
|
||||
DualIntrusiveAllocatorThreadPool DualListData;
|
||||
|
||||
OrderedNode* InvalidNode;
|
||||
OrderedNode* CurrentCodeBlock {};
|
||||
fextl::vector<OrderedNode*> CodeBlocks;
|
||||
Ref InvalidNode;
|
||||
Ref CurrentCodeBlock {};
|
||||
fextl::vector<Ref> CodeBlocks;
|
||||
uint64_t Entry;
|
||||
};
|
||||
|
||||
|
||||
@@ -1,724 +0,0 @@
|
||||
// SPDX-License-Identifier: MIT
|
||||
/*
|
||||
$info$
|
||||
meta: ir|parser ~ Text -> IR
|
||||
tags: ir|parser
|
||||
$end_info$
|
||||
*/
|
||||
|
||||
#include "Interface/IR/IREmitter.h"
|
||||
|
||||
#include <FEXCore/IR/IR.h>
|
||||
#include <FEXCore/Utils/LogManager.h>
|
||||
#include <FEXCore/Utils/StringUtils.h>
|
||||
#include <FEXCore/fextl/sstream.h>
|
||||
#include <FEXCore/fextl/string.h>
|
||||
#include <FEXCore/fextl/unordered_map.h>
|
||||
#include <FEXCore/fextl/vector.h>
|
||||
|
||||
#include <algorithm>
|
||||
#include <array>
|
||||
#include <cstdint>
|
||||
#include <errno.h>
|
||||
#include <memory>
|
||||
#include <stdio.h>
|
||||
#include <stdlib.h>
|
||||
#include <string_view>
|
||||
#include <utility>
|
||||
#include <istream>
|
||||
#include <unordered_map>
|
||||
|
||||
namespace FEXCore::IR {
|
||||
namespace {
|
||||
|
||||
enum class DecodeFailure {
|
||||
DECODE_OKAY,
|
||||
DECODE_UNKNOWN_TYPE,
|
||||
DECODE_INVALID,
|
||||
DECODE_INVALIDCHAR,
|
||||
DECODE_INVALIDRANGE,
|
||||
DECODE_INVALIDREGISTERCLASS,
|
||||
DECODE_UNKNOWN_SSA,
|
||||
DECODE_INVALID_CONDFLAG,
|
||||
DECODE_INVALID_MEMOFFSETTYPE,
|
||||
DECODE_INVALID_FENCETYPE,
|
||||
DECODE_INVALID_BREAKTYPE,
|
||||
DECODE_INVALID_OPSIZE,
|
||||
};
|
||||
|
||||
fextl::string DecodeErrorToString(DecodeFailure Failure) {
|
||||
switch (Failure) {
|
||||
case DecodeFailure::DECODE_OKAY: return "Okay";
|
||||
case DecodeFailure::DECODE_UNKNOWN_TYPE: return "Unknown Type";
|
||||
case DecodeFailure::DECODE_INVALID: return "Invalid";
|
||||
case DecodeFailure::DECODE_INVALIDCHAR: return "Invalid starting char";
|
||||
case DecodeFailure::DECODE_INVALIDRANGE: return "Invalid integer range";
|
||||
case DecodeFailure::DECODE_INVALIDREGISTERCLASS: return "Invalid register class";
|
||||
case DecodeFailure::DECODE_UNKNOWN_SSA: return "Unknown SSA value";
|
||||
case DecodeFailure::DECODE_INVALID_CONDFLAG: return "Invalid Conditional name";
|
||||
case DecodeFailure::DECODE_INVALID_MEMOFFSETTYPE: return "Invalid Memory Offset Type";
|
||||
case DecodeFailure::DECODE_INVALID_FENCETYPE: return "Invalid Fence Type";
|
||||
case DecodeFailure::DECODE_INVALID_BREAKTYPE: return "Invalid Break Reason Type";
|
||||
case DecodeFailure::DECODE_INVALID_OPSIZE: return "Invalid Operation size name";
|
||||
}
|
||||
return "Unknown Error";
|
||||
}
|
||||
|
||||
class IRParser : public FEXCore::IR::IREmitter {
|
||||
public:
|
||||
template<typename Type>
|
||||
std::pair<DecodeFailure, Type> DecodeValue(const fextl::string& Arg) {
|
||||
return {DecodeFailure::DECODE_UNKNOWN_TYPE, {}};
|
||||
}
|
||||
|
||||
template<>
|
||||
std::pair<DecodeFailure, uint8_t> DecodeValue(const fextl::string& Arg) {
|
||||
if (Arg.at(0) != '#') {
|
||||
return {DecodeFailure::DECODE_INVALIDCHAR, 0};
|
||||
}
|
||||
|
||||
uint8_t Result = strtoul(&Arg.at(1), nullptr, 0);
|
||||
if (errno == ERANGE) {
|
||||
return {DecodeFailure::DECODE_INVALIDRANGE, 0};
|
||||
}
|
||||
return {DecodeFailure::DECODE_OKAY, Result};
|
||||
}
|
||||
|
||||
template<>
|
||||
std::pair<DecodeFailure, bool> DecodeValue(const fextl::string& Arg) {
|
||||
if (Arg.at(0) != '#') {
|
||||
return {DecodeFailure::DECODE_INVALIDCHAR, 0};
|
||||
}
|
||||
|
||||
uint8_t Result = strtoul(&Arg.at(1), nullptr, 0);
|
||||
if (errno == ERANGE || Result > 1) {
|
||||
return {DecodeFailure::DECODE_INVALIDRANGE, 0};
|
||||
}
|
||||
return {DecodeFailure::DECODE_OKAY, Result != 0};
|
||||
}
|
||||
|
||||
template<>
|
||||
std::pair<DecodeFailure, uint16_t> DecodeValue(const fextl::string& Arg) {
|
||||
if (Arg.at(0) != '#') {
|
||||
return {DecodeFailure::DECODE_INVALIDCHAR, 0};
|
||||
}
|
||||
|
||||
uint16_t Result = strtoul(&Arg.at(1), nullptr, 0);
|
||||
if (errno == ERANGE) {
|
||||
return {DecodeFailure::DECODE_INVALIDRANGE, 0};
|
||||
}
|
||||
return {DecodeFailure::DECODE_OKAY, Result};
|
||||
}
|
||||
|
||||
template<>
|
||||
std::pair<DecodeFailure, uint32_t> DecodeValue(const fextl::string& Arg) {
|
||||
if (Arg.at(0) != '#') {
|
||||
return {DecodeFailure::DECODE_INVALIDCHAR, 0};
|
||||
}
|
||||
|
||||
uint32_t Result = strtoul(&Arg.at(1), nullptr, 0);
|
||||
if (errno == ERANGE) {
|
||||
return {DecodeFailure::DECODE_INVALIDRANGE, 0};
|
||||
}
|
||||
return {DecodeFailure::DECODE_OKAY, Result};
|
||||
}
|
||||
|
||||
template<>
|
||||
std::pair<DecodeFailure, uint64_t> DecodeValue(const fextl::string& Arg) {
|
||||
if (Arg.at(0) != '#') {
|
||||
return {DecodeFailure::DECODE_INVALIDCHAR, 0};
|
||||
}
|
||||
|
||||
uint64_t Result = strtoull(&Arg.at(1), nullptr, 0);
|
||||
if (errno == ERANGE) {
|
||||
return {DecodeFailure::DECODE_INVALIDRANGE, 0};
|
||||
}
|
||||
return {DecodeFailure::DECODE_OKAY, Result};
|
||||
}
|
||||
|
||||
template<>
|
||||
std::pair<DecodeFailure, int64_t> DecodeValue(const fextl::string& Arg) {
|
||||
if (Arg.at(0) != '#') {
|
||||
return {DecodeFailure::DECODE_INVALIDCHAR, 0};
|
||||
}
|
||||
|
||||
int64_t Result = (int64_t)strtoull(&Arg.at(1), nullptr, 0);
|
||||
if (errno == ERANGE) {
|
||||
return {DecodeFailure::DECODE_INVALIDRANGE, 0};
|
||||
}
|
||||
return {DecodeFailure::DECODE_OKAY, Result};
|
||||
}
|
||||
|
||||
template<>
|
||||
std::pair<DecodeFailure, IR::SHA256Sum> DecodeValue(const fextl::string& Arg) {
|
||||
IR::SHA256Sum Result;
|
||||
|
||||
if (Arg.at(0) != 's' || Arg.at(1) != 'h' || Arg.at(2) != 'a' || Arg.at(3) != '2' || Arg.at(4) != '5' || Arg.at(5) != '6' || Arg.at(6) != ':') {
|
||||
return {DecodeFailure::DECODE_INVALIDCHAR, Result};
|
||||
}
|
||||
|
||||
auto GetDigit = [](const fextl::string& Arg, int pos, uint8_t* val) {
|
||||
auto chr = Arg.at(pos);
|
||||
if (chr >= '0' && chr <= '9') {
|
||||
*val = chr - '0';
|
||||
return true;
|
||||
} else if (chr >= 'a' && chr <= 'f') {
|
||||
*val = 10 + chr - 'a';
|
||||
return true;
|
||||
} else {
|
||||
return false;
|
||||
}
|
||||
};
|
||||
|
||||
for (size_t i = 0; i < sizeof(Result.data); i++) {
|
||||
uint8_t high, low;
|
||||
if (!GetDigit(Arg, 7 + 2 * i + 0, &high) || !GetDigit(Arg, 7 + 2 * i + 1, &low)) {
|
||||
return {DecodeFailure::DECODE_INVALIDRANGE, Result};
|
||||
}
|
||||
Result.data[i] = high * 16 + low;
|
||||
}
|
||||
|
||||
return {DecodeFailure::DECODE_OKAY, Result};
|
||||
}
|
||||
|
||||
template<>
|
||||
std::pair<DecodeFailure, FEXCore::IR::RegisterClassType> DecodeValue(const fextl::string& Arg) {
|
||||
if (Arg == "GPR") {
|
||||
return {DecodeFailure::DECODE_OKAY, FEXCore::IR::GPRClass};
|
||||
} else if (Arg == "FPR") {
|
||||
return {DecodeFailure::DECODE_OKAY, FEXCore::IR::FPRClass};
|
||||
} else if (Arg == "GPRFixed") {
|
||||
return {DecodeFailure::DECODE_OKAY, FEXCore::IR::GPRFixedClass};
|
||||
} else if (Arg == "FPRFixed") {
|
||||
return {DecodeFailure::DECODE_OKAY, FEXCore::IR::FPRFixedClass};
|
||||
} else if (Arg == "GPRPair") {
|
||||
return {DecodeFailure::DECODE_OKAY, FEXCore::IR::GPRPairClass};
|
||||
} else if (Arg == "Complex") {
|
||||
return {DecodeFailure::DECODE_OKAY, FEXCore::IR::ComplexClass};
|
||||
}
|
||||
|
||||
return {DecodeFailure::DECODE_INVALIDREGISTERCLASS, FEXCore::IR::InvalidClass};
|
||||
}
|
||||
|
||||
template<>
|
||||
std::pair<DecodeFailure, FEXCore::IR::TypeDefinition> DecodeValue(const fextl::string& Arg) {
|
||||
uint8_t Size {}, Elements {1};
|
||||
int NumArgs = sscanf(Arg.c_str(), "i%hhdv%hhd", &Size, &Elements);
|
||||
|
||||
if (NumArgs != 1 && NumArgs != 2) {
|
||||
return {DecodeFailure::DECODE_INVALID, {}};
|
||||
}
|
||||
|
||||
return {DecodeFailure::DECODE_OKAY, FEXCore::IR::TypeDefinition::Create(Size / 8, Elements)};
|
||||
}
|
||||
|
||||
template<>
|
||||
std::pair<DecodeFailure, FEXCore::IR::CondClassType> DecodeValue(const fextl::string& Arg) {
|
||||
static constexpr std::array<std::string_view, 22> CondNames = {"EQ", "NEQ", "UGE", "ULT", "MI", "PL", "VS", "VC",
|
||||
"UGT", "ULE", "SGE", "SLT", "SGT", "SLE", "ANDZ", "ANDNZ",
|
||||
"FLU", "FGE", "FLEU", "FGT", "FU", "FNU"};
|
||||
|
||||
for (size_t i = 0; i < CondNames.size(); ++i) {
|
||||
if (CondNames[i] == Arg) {
|
||||
return {DecodeFailure::DECODE_OKAY, CondClassType {static_cast<uint8_t>(i)}};
|
||||
}
|
||||
}
|
||||
return {DecodeFailure::DECODE_INVALID_CONDFLAG, {}};
|
||||
}
|
||||
|
||||
template<>
|
||||
std::pair<DecodeFailure, FEXCore::IR::MemOffsetType> DecodeValue(const fextl::string& Arg) {
|
||||
static constexpr std::array<std::string_view, 3> Names = {
|
||||
"SXTX",
|
||||
"UXTW",
|
||||
"SXTW",
|
||||
};
|
||||
|
||||
for (size_t i = 0; i < Names.size(); ++i) {
|
||||
if (Names[i] == Arg) {
|
||||
return {DecodeFailure::DECODE_OKAY, MemOffsetType {static_cast<uint8_t>(i)}};
|
||||
}
|
||||
}
|
||||
return {DecodeFailure::DECODE_INVALID_MEMOFFSETTYPE, {}};
|
||||
}
|
||||
|
||||
template<>
|
||||
std::pair<DecodeFailure, FEXCore::IR::FenceType> DecodeValue(const fextl::string& Arg) {
|
||||
static constexpr std::array<std::string_view, 3> Names = {
|
||||
"Loads",
|
||||
"Stores",
|
||||
"LoadStores",
|
||||
};
|
||||
|
||||
for (size_t i = 0; i < Names.size(); ++i) {
|
||||
if (Names[i] == Arg) {
|
||||
return {DecodeFailure::DECODE_OKAY, FenceType {static_cast<uint8_t>(i)}};
|
||||
}
|
||||
}
|
||||
return {DecodeFailure::DECODE_INVALID_FENCETYPE, {}};
|
||||
}
|
||||
|
||||
template<>
|
||||
std::pair<DecodeFailure, FEXCore::IR::BreakDefinition> DecodeValue(const fextl::string& Arg) {
|
||||
uint32_t tmp {};
|
||||
fextl::stringstream ss {Arg};
|
||||
BreakDefinition Reason {};
|
||||
|
||||
// Seek past '{'
|
||||
ss.seekg(1, std::ios::cur);
|
||||
ss >> Reason.ErrorRegister;
|
||||
|
||||
// Seek past '.'
|
||||
ss.seekg(1, std::ios::cur);
|
||||
ss >> tmp;
|
||||
Reason.Signal = tmp;
|
||||
|
||||
// Seek past '.'
|
||||
ss.seekg(1, std::ios::cur);
|
||||
ss >> tmp;
|
||||
Reason.TrapNumber = tmp;
|
||||
|
||||
// Seek past '.'
|
||||
ss.seekg(1, std::ios::cur);
|
||||
ss >> tmp;
|
||||
Reason.si_code = tmp;
|
||||
|
||||
if (ss.fail()) {
|
||||
return {DecodeFailure::DECODE_INVALIDCHAR, {}};
|
||||
} else {
|
||||
return {DecodeFailure::DECODE_OKAY, Reason};
|
||||
}
|
||||
}
|
||||
|
||||
template<>
|
||||
std::pair<DecodeFailure, FEXCore::IR::OpSize> DecodeValue(const fextl::string& Arg) {
|
||||
static constexpr std::array<std::pair<std::string_view, FEXCore::IR::OpSize>, 6> Names = {{
|
||||
{"i8", OpSize::i8Bit},
|
||||
{"i16", OpSize::i16Bit},
|
||||
{"i32", OpSize::i32Bit},
|
||||
{"i64", OpSize::i64Bit},
|
||||
{"i128", OpSize::i128Bit},
|
||||
{"i256", OpSize::i256Bit},
|
||||
}};
|
||||
|
||||
for (size_t i = 0; i < Names.size(); ++i) {
|
||||
if (Names[i].first == Arg) {
|
||||
return {DecodeFailure::DECODE_OKAY, Names[i].second};
|
||||
}
|
||||
}
|
||||
return {DecodeFailure::DECODE_INVALID_OPSIZE, {}};
|
||||
}
|
||||
|
||||
template<>
|
||||
std::pair<DecodeFailure, OrderedNode*> DecodeValue(const fextl::string& Arg) {
|
||||
if (Arg.at(0) != '%') {
|
||||
return {DecodeFailure::DECODE_INVALIDCHAR, 0};
|
||||
}
|
||||
|
||||
// Strip off the type qualifier from the ssa value
|
||||
fextl::string SSAName = FEXCore::StringUtils::Trim(Arg);
|
||||
const size_t ArgEnd = SSAName.find_first_of(' ');
|
||||
|
||||
if (ArgEnd != fextl::string::npos) {
|
||||
SSAName = SSAName.substr(0, ArgEnd);
|
||||
}
|
||||
|
||||
// Forward declarations may make this not succed
|
||||
auto Op = SSANameMapper.find(SSAName);
|
||||
if (Op == SSANameMapper.end()) {
|
||||
return {DecodeFailure::DECODE_UNKNOWN_SSA, nullptr};
|
||||
}
|
||||
|
||||
return {DecodeFailure::DECODE_OKAY, Op->second};
|
||||
}
|
||||
|
||||
struct LineDefinition {
|
||||
size_t LineNumber;
|
||||
bool HasDefinition {};
|
||||
fextl::string Definition {};
|
||||
FEXCore::IR::TypeDefinition Size {};
|
||||
fextl::string IROp {};
|
||||
FEXCore::IR::IROps OpEnum;
|
||||
bool HasArgs {};
|
||||
fextl::vector<fextl::string> Args;
|
||||
OrderedNode* Node {};
|
||||
};
|
||||
|
||||
fextl::vector<fextl::string> Lines;
|
||||
fextl::unordered_map<fextl::string, OrderedNode*> SSANameMapper;
|
||||
fextl::vector<LineDefinition> Defs;
|
||||
LineDefinition* CurrentDef {};
|
||||
fextl::unordered_map<std::string_view, FEXCore::IR::IROps> NameToOpMap;
|
||||
|
||||
IRParser(FEXCore::Utils::IntrusivePooledAllocator& ThreadAllocator, fextl::stringstream& MapsStream)
|
||||
: IREmitter {ThreadAllocator} {
|
||||
InitializeNameMap();
|
||||
|
||||
fextl::string Line;
|
||||
while (std::getline(MapsStream, Line)) {
|
||||
if (MapsStream.eof()) {
|
||||
break;
|
||||
}
|
||||
if (MapsStream.fail()) {
|
||||
LogMan::Msg::EFmt("Failed to getline on line: {}", Lines.size());
|
||||
return;
|
||||
}
|
||||
Lines.emplace_back(Line);
|
||||
}
|
||||
|
||||
ResetWorkingList();
|
||||
Loaded = Parse();
|
||||
}
|
||||
|
||||
bool Loaded = false;
|
||||
|
||||
bool Parse() {
|
||||
const auto CheckPrintError = [&](const LineDefinition& Def, DecodeFailure Failure) -> bool {
|
||||
if (Failure != DecodeFailure::DECODE_OKAY) {
|
||||
LogMan::Msg::EFmt("Error on Line: {}", Def.LineNumber);
|
||||
LogMan::Msg::EFmt("{}", Lines[Def.LineNumber]);
|
||||
LogMan::Msg::EFmt("Value Couldn't be decoded due to {}", DecodeErrorToString(Failure));
|
||||
return false;
|
||||
}
|
||||
|
||||
return true;
|
||||
};
|
||||
|
||||
const auto CheckPrintErrorArg = [&](const LineDefinition& Def, DecodeFailure Failure, size_t Arg) -> bool {
|
||||
if (Failure != DecodeFailure::DECODE_OKAY) {
|
||||
LogMan::Msg::EFmt("Error on Line: {}", Def.LineNumber);
|
||||
LogMan::Msg::EFmt("{}", Lines[Def.LineNumber]);
|
||||
LogMan::Msg::EFmt("Argument Number {}: {}", Arg + 1, Def.Args[Arg]);
|
||||
LogMan::Msg::EFmt("Value Couldn't be decoded due to {}", DecodeErrorToString(Failure));
|
||||
return false;
|
||||
}
|
||||
|
||||
return true;
|
||||
};
|
||||
|
||||
// String parse every line for our definitions
|
||||
for (size_t i = 0; i < Lines.size(); ++i) {
|
||||
fextl::string Line = Lines[i];
|
||||
LineDefinition Def {};
|
||||
CurrentDef = &Def;
|
||||
Def.LineNumber = i;
|
||||
|
||||
Line = FEXCore::StringUtils::Trim(Line);
|
||||
|
||||
// Skip empty lines
|
||||
if (Line.empty()) {
|
||||
continue;
|
||||
}
|
||||
|
||||
if (Line[0] == ';') {
|
||||
// This is a comment line
|
||||
// Skip it
|
||||
continue;
|
||||
}
|
||||
|
||||
size_t CurrentPos {};
|
||||
// Let's see if this node is assigning something first
|
||||
if (Line[0] == '%') {
|
||||
size_t DefinitionEnd = fextl::string::npos;
|
||||
if ((DefinitionEnd = Line.find_first_of('=', CurrentPos)) != fextl::string::npos) {
|
||||
Def.Definition = Line.substr(0, DefinitionEnd);
|
||||
Def.Definition = FEXCore::StringUtils::Trim(Def.Definition);
|
||||
Def.HasDefinition = true;
|
||||
CurrentPos = DefinitionEnd + 1; // +1 to ensure we go past then assignment
|
||||
} else {
|
||||
LogMan::Msg::EFmt("Error on Line: {}", i);
|
||||
LogMan::Msg::EFmt("{}", Lines[i]);
|
||||
LogMan::Msg::EFmt("SSA declaration without assignment");
|
||||
return false;
|
||||
}
|
||||
}
|
||||
|
||||
// Check if we are pulling in some IR from the IR Printer
|
||||
// Prints (%%d) at the start of lines without a definition
|
||||
if (Line[0] == '(') {
|
||||
size_t DefinitionEnd = fextl::string::npos;
|
||||
if ((DefinitionEnd = Line.find_first_of(')', CurrentPos)) != fextl::string::npos) {
|
||||
size_t SSAEnd = fextl::string::npos;
|
||||
if ((SSAEnd = Line.find_last_of(' ', DefinitionEnd)) != fextl::string::npos) {
|
||||
fextl::string Type = Line.substr(SSAEnd + 1, DefinitionEnd - SSAEnd - 1);
|
||||
Type = FEXCore::StringUtils::Trim(Type);
|
||||
|
||||
auto DefinitionSize = DecodeValue<FEXCore::IR::TypeDefinition>(Type);
|
||||
if (!CheckPrintError(Def, DefinitionSize.first)) {
|
||||
return false;
|
||||
}
|
||||
Def.Size = DefinitionSize.second;
|
||||
}
|
||||
|
||||
Def.Definition = FEXCore::StringUtils::Trim(Line.substr(1, std::min(DefinitionEnd, SSAEnd) - 1));
|
||||
|
||||
CurrentPos = DefinitionEnd + 1;
|
||||
} else {
|
||||
LogMan::Msg::EFmt("Error on Line: {}", i);
|
||||
LogMan::Msg::EFmt("{}", Lines[i]);
|
||||
LogMan::Msg::EFmt("SSA value with numbered SSA provided but no closing parentheses");
|
||||
return false;
|
||||
}
|
||||
}
|
||||
|
||||
if (Def.HasDefinition) {
|
||||
// Let's check if we have a size declared with this variable
|
||||
size_t NameEnd = fextl::string::npos;
|
||||
if ((NameEnd = Def.Definition.find_first_of(' ')) != fextl::string::npos) {
|
||||
fextl::string Type = Def.Definition.substr(NameEnd + 1);
|
||||
Type = FEXCore::StringUtils::Trim(Type);
|
||||
Def.Definition = FEXCore::StringUtils::Trim(Def.Definition.substr(0, NameEnd));
|
||||
|
||||
auto DefinitionSize = DecodeValue<FEXCore::IR::TypeDefinition>(Type);
|
||||
if (!CheckPrintError(Def, DefinitionSize.first)) {
|
||||
return false;
|
||||
}
|
||||
Def.Size = DefinitionSize.second;
|
||||
}
|
||||
|
||||
if (Def.Definition == "%Invalid") {
|
||||
LogMan::Msg::EFmt("Error on Line: {}", i);
|
||||
LogMan::Msg::EFmt("{}", Lines[i]);
|
||||
LogMan::Msg::EFmt("Definition tried to define reserved %Invalid ssa node");
|
||||
return false;
|
||||
}
|
||||
}
|
||||
|
||||
// Let's get the IR op
|
||||
size_t OpNameEnd = fextl::string::npos;
|
||||
fextl::string RemainingLine = FEXCore::StringUtils::Trim(Line.substr(CurrentPos));
|
||||
CurrentPos = 0;
|
||||
if ((OpNameEnd = RemainingLine.find_first_of(" \t\n\r\0", CurrentPos)) != fextl::string::npos) {
|
||||
Def.IROp = RemainingLine.substr(CurrentPos, OpNameEnd);
|
||||
Def.IROp = FEXCore::StringUtils::Trim(Def.IROp);
|
||||
Def.HasArgs = true;
|
||||
CurrentPos = OpNameEnd;
|
||||
} else {
|
||||
if (RemainingLine.empty()) {
|
||||
LogMan::Msg::EFmt("Error on Line: {}", i);
|
||||
LogMan::Msg::EFmt("{}", Lines[i]);
|
||||
LogMan::Msg::EFmt("Line without an IROp?");
|
||||
return false;
|
||||
}
|
||||
|
||||
Def.IROp = RemainingLine;
|
||||
Def.HasArgs = false;
|
||||
}
|
||||
|
||||
if (Def.HasArgs) {
|
||||
RemainingLine = FEXCore::StringUtils::Trim(RemainingLine.substr(CurrentPos));
|
||||
if (RemainingLine.empty()) {
|
||||
// How did we get here?
|
||||
Def.HasArgs = false;
|
||||
} else {
|
||||
while (!RemainingLine.empty()) {
|
||||
const size_t ArgEnd = RemainingLine.find(',');
|
||||
fextl::string Arg = FEXCore::StringUtils::Trim(RemainingLine.substr(0, ArgEnd));
|
||||
|
||||
Def.Args.emplace_back(std::move(Arg));
|
||||
|
||||
RemainingLine.erase(0, ArgEnd + 1); // +1 to ensure we go past the ','
|
||||
if (ArgEnd == fextl::string::npos) {
|
||||
break;
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
CurrentDef = &Defs.emplace_back(std::move(Def));
|
||||
}
|
||||
|
||||
// Ensure all of the ops are real ops
|
||||
for (size_t i = 0; i < Defs.size(); ++i) {
|
||||
auto& Def = Defs[i];
|
||||
auto Op = NameToOpMap.find(Def.IROp);
|
||||
if (Op == NameToOpMap.end()) {
|
||||
LogMan::Msg::EFmt("Error on Line: {}", Def.LineNumber);
|
||||
LogMan::Msg::EFmt("{}", Lines[Def.LineNumber]);
|
||||
LogMan::Msg::EFmt("IROp '{}' doesn't exist", Def.IROp);
|
||||
return false;
|
||||
}
|
||||
Def.OpEnum = Op->second;
|
||||
}
|
||||
|
||||
// Emit the header op
|
||||
IRPair<IROp_IRHeader> IRHeader;
|
||||
{
|
||||
auto& Def = Defs[0];
|
||||
CurrentDef = &Def;
|
||||
if (Def.OpEnum != FEXCore::IR::IROps::OP_IRHEADER) {
|
||||
LogMan::Msg::EFmt("Error on Line: {}", Def.LineNumber);
|
||||
LogMan::Msg::EFmt("{}", Lines[Def.LineNumber]);
|
||||
LogMan::Msg::EFmt("First op needs to be IRHeader. Was '{}'", Def.IROp);
|
||||
return false;
|
||||
}
|
||||
|
||||
auto OriginalRIP = DecodeValue<uint64_t>(Def.Args[1]);
|
||||
|
||||
if (!CheckPrintError(Def, OriginalRIP.first)) {
|
||||
return false;
|
||||
}
|
||||
|
||||
auto CodeBlockCount = DecodeValue<uint64_t>(Def.Args[2]);
|
||||
|
||||
if (!CheckPrintError(Def, CodeBlockCount.first)) {
|
||||
return false;
|
||||
}
|
||||
|
||||
auto InstructionCount = DecodeValue<uint64_t>(Def.Args[3]);
|
||||
|
||||
if (!CheckPrintError(Def, InstructionCount.first)) {
|
||||
return false;
|
||||
}
|
||||
|
||||
IRHeader = _IRHeader(InvalidNode, OriginalRIP.second, CodeBlockCount.second, InstructionCount.second);
|
||||
}
|
||||
|
||||
SetWriteCursor(nullptr); // isolate the header from everything following
|
||||
|
||||
// Initialize SSANameMapper with Invalid value
|
||||
SSANameMapper.insert_or_assign("%Invalid", Invalid());
|
||||
|
||||
// Spin through the blocks and generate basic block ops
|
||||
for (size_t i = 0; i < Defs.size(); ++i) {
|
||||
auto& Def = Defs[i];
|
||||
if (Def.OpEnum == FEXCore::IR::IROps::OP_CODEBLOCK) {
|
||||
auto CodeBlock = _CodeBlock(InvalidNode, InvalidNode);
|
||||
SSANameMapper.insert_or_assign(Def.Definition, CodeBlock.Node);
|
||||
Def.Node = CodeBlock.Node;
|
||||
|
||||
if (i == 1) {
|
||||
// First code block is the entry block
|
||||
// Link the header to the first block
|
||||
IRHeader.first->Blocks = CodeBlock.Node->Wrapped(DualListData.ListBegin());
|
||||
}
|
||||
CodeBlocks.emplace_back(CodeBlock.Node);
|
||||
}
|
||||
}
|
||||
SetWriteCursor(nullptr); // isolate the block headers too
|
||||
|
||||
// Spin through all the definitions and add the ops to the basic blocks
|
||||
OrderedNode* CurrentBlock {};
|
||||
FEXCore::IR::IROp_CodeBlock* CurrentBlockOp {};
|
||||
for (size_t i = 1; i < Defs.size(); ++i) {
|
||||
auto& Def = Defs[i];
|
||||
CurrentDef = &Def;
|
||||
|
||||
switch (Def.OpEnum) {
|
||||
// Special handled
|
||||
case FEXCore::IR::IROps::OP_IRHEADER:
|
||||
LogMan::Msg::EFmt("Error on Line: {}", Def.LineNumber);
|
||||
LogMan::Msg::EFmt("{}", Lines[Def.LineNumber]);
|
||||
LogMan::Msg::EFmt("IRHEADER used in the middle of the block!");
|
||||
return false; // only one OP_IRHEADER allowed per block
|
||||
|
||||
case FEXCore::IR::IROps::OP_CODEBLOCK: {
|
||||
SetWriteCursor(nullptr); // isolate from previous block
|
||||
if (CurrentBlock != nullptr) {
|
||||
LogMan::Msg::EFmt("Error on Line: {}", Def.LineNumber);
|
||||
LogMan::Msg::EFmt("{}", Lines[Def.LineNumber]);
|
||||
LogMan::Msg::EFmt("CodeBlock being used inside of already existing codeblock!");
|
||||
return false;
|
||||
}
|
||||
|
||||
CurrentBlock = Def.Node;
|
||||
CurrentBlockOp = CurrentBlock->Op(DualListData.DataBegin())->CW<FEXCore::IR::IROp_CodeBlock>();
|
||||
break;
|
||||
}
|
||||
|
||||
|
||||
case FEXCore::IR::IROps::OP_BEGINBLOCK: {
|
||||
if (CurrentBlock == nullptr) {
|
||||
LogMan::Msg::EFmt("Error on Line: {}", Def.LineNumber);
|
||||
LogMan::Msg::EFmt("{}", Lines[Def.LineNumber]);
|
||||
LogMan::Msg::EFmt("EndBlock being used outside of a block!");
|
||||
return false;
|
||||
}
|
||||
|
||||
auto Adjust = DecodeValue<OrderedNode*>(Def.Args[0]);
|
||||
if (!CheckPrintError(Def, Adjust.first)) {
|
||||
return false;
|
||||
}
|
||||
|
||||
Def.Node = _BeginBlock(Adjust.second);
|
||||
CurrentBlockOp->Begin = Def.Node->Wrapped(DualListData.ListBegin());
|
||||
break;
|
||||
}
|
||||
|
||||
case FEXCore::IR::IROps::OP_ENDBLOCK: {
|
||||
if (CurrentBlock == nullptr) {
|
||||
LogMan::Msg::EFmt("Error on Line: {}", Def.LineNumber);
|
||||
LogMan::Msg::EFmt("{}", Lines[Def.LineNumber]);
|
||||
LogMan::Msg::EFmt("EndBlock being used outside of a block!");
|
||||
return false;
|
||||
}
|
||||
|
||||
auto Adjust = DecodeValue<OrderedNode*>(Def.Args[0]);
|
||||
if (!CheckPrintError(Def, Adjust.first)) {
|
||||
return false;
|
||||
}
|
||||
|
||||
Def.Node = _EndBlock(Adjust.second);
|
||||
CurrentBlockOp->Last = Def.Node->Wrapped(DualListData.ListBegin());
|
||||
|
||||
CurrentBlock = nullptr;
|
||||
CurrentBlockOp = nullptr;
|
||||
|
||||
break;
|
||||
}
|
||||
|
||||
case FEXCore::IR::IROps::OP_DUMMY: {
|
||||
LogMan::Msg::EFmt("Error on Line: {}", Def.LineNumber);
|
||||
LogMan::Msg::EFmt("{}", Lines[Def.LineNumber]);
|
||||
LogMan::Msg::EFmt("Dummy op must not be used");
|
||||
break;
|
||||
}
|
||||
#define IROP_PARSER_SWITCH_HELPERS
|
||||
#include <FEXCore/IR/IRDefines.inc>
|
||||
default: {
|
||||
LogMan::Msg::EFmt("Error on Line: {}", Def.LineNumber);
|
||||
LogMan::Msg::EFmt("{}", Lines[Def.LineNumber]);
|
||||
LogMan::Msg::EFmt("Unhandled Op enum '{}' in parser", Def.IROp);
|
||||
return false;
|
||||
}
|
||||
}
|
||||
|
||||
if (Def.HasDefinition) {
|
||||
auto IROp = Def.Node->Op(DualListData.DataBegin());
|
||||
if (Def.Size.Elements()) {
|
||||
IROp->Size = Def.Size.Bytes() * Def.Size.Elements();
|
||||
IROp->ElementSize = Def.Size.Bytes();
|
||||
} else {
|
||||
IROp->Size = Def.Size.Bytes();
|
||||
IROp->ElementSize = 0;
|
||||
}
|
||||
SSANameMapper.insert_or_assign(Def.Definition, Def.Node);
|
||||
}
|
||||
}
|
||||
|
||||
return true;
|
||||
}
|
||||
|
||||
void InitializeNameMap() {
|
||||
if (NameToOpMap.empty()) {
|
||||
for (FEXCore::IR::IROps Op = FEXCore::IR::IROps::OP_DUMMY; Op <= FEXCore::IR::IROps::OP_LAST;
|
||||
Op = static_cast<FEXCore::IR::IROps>(static_cast<uint32_t>(Op) + 1)) {
|
||||
NameToOpMap.insert_or_assign(FEXCore::IR::GetName(Op), Op);
|
||||
}
|
||||
}
|
||||
}
|
||||
};
|
||||
|
||||
} // namespace
|
||||
|
||||
fextl::unique_ptr<IREmitter> Parse(FEXCore::Utils::IntrusivePooledAllocator& ThreadAllocator, fextl::stringstream& MapsStream) {
|
||||
auto parser = fextl::make_unique<IRParser>(ThreadAllocator, MapsStream);
|
||||
|
||||
if (parser->Loaded) {
|
||||
return parser;
|
||||
} else {
|
||||
return nullptr;
|
||||
}
|
||||
}
|
||||
|
||||
} // namespace FEXCore::IR
|
||||
@@ -144,54 +144,22 @@ private:
|
||||
Utils::FixedSizePooledAllocation<uintptr_t, 5000, 500> PoolObject;
|
||||
};
|
||||
|
||||
class IRListView final : public FEXCore::Allocator::FEXAllocOperators {
|
||||
enum Flags {
|
||||
FLAG_IsCopy = 1,
|
||||
FLAG_Shared = 2,
|
||||
};
|
||||
|
||||
class IRListView final {
|
||||
public:
|
||||
IRListView() = delete;
|
||||
IRListView(IRListView&&) = delete;
|
||||
|
||||
IRListView(DualIntrusiveAllocator* Data, bool _IsCopy) {
|
||||
SetCopy(_IsCopy);
|
||||
DataSize = Data->DataSize();
|
||||
ListSize = Data->ListSize();
|
||||
IRListView(DualIntrusiveAllocator* Data)
|
||||
: IRListView(reinterpret_cast<void*>(Data->DataBegin()), reinterpret_cast<void*>(Data->ListBegin()), Data->DataSize(), Data->ListSize()) {}
|
||||
|
||||
if (_IsCopy) {
|
||||
IRDataInternal = FEXCore::Allocator::malloc(DataSize + ListSize);
|
||||
ListDataInternal = reinterpret_cast<void*>(reinterpret_cast<uintptr_t>(IRDataInternal) + DataSize);
|
||||
memcpy(IRDataInternal, reinterpret_cast<void*>(Data->DataBegin()), DataSize);
|
||||
memcpy(ListDataInternal, reinterpret_cast<void*>(Data->ListBegin()), ListSize);
|
||||
} else {
|
||||
// We are just pointing to the data
|
||||
IRDataInternal = reinterpret_cast<void*>(Data->DataBegin());
|
||||
ListDataInternal = reinterpret_cast<void*>(Data->ListBegin());
|
||||
}
|
||||
}
|
||||
IRListView(IRListView* Old)
|
||||
: IRListView(Old->IRDataInternal, Old->ListDataInternal, Old->DataSize, Old->ListSize) {}
|
||||
|
||||
IRListView(IRListView* Old, bool _IsCopy) {
|
||||
SetCopy(_IsCopy);
|
||||
DataSize = Old->DataSize;
|
||||
ListSize = Old->ListSize;
|
||||
if (_IsCopy) {
|
||||
IRDataInternal = FEXCore::Allocator::malloc(DataSize + ListSize);
|
||||
ListDataInternal = reinterpret_cast<void*>(reinterpret_cast<uintptr_t>(IRDataInternal) + DataSize);
|
||||
memcpy(IRDataInternal, Old->IRDataInternal, DataSize);
|
||||
memcpy(ListDataInternal, Old->ListDataInternal, ListSize);
|
||||
} else {
|
||||
IRDataInternal = Old->IRDataInternal;
|
||||
ListDataInternal = Old->ListDataInternal;
|
||||
}
|
||||
}
|
||||
|
||||
~IRListView() {
|
||||
if (IsCopy()) {
|
||||
FEXCore::Allocator::free(IRDataInternal);
|
||||
// ListData is just offset from IRData
|
||||
}
|
||||
}
|
||||
IRListView(void* IRData_, void* ListData_, size_t DataSize_, size_t ListSize_)
|
||||
: IRDataInternal(IRData_)
|
||||
, ListDataInternal(ListData_)
|
||||
, DataSize(DataSize_)
|
||||
, ListSize(ListSize_) {}
|
||||
|
||||
void Serialize(FEXCore::Context::AOTIRWriter& stream) const {
|
||||
void* nul = nullptr;
|
||||
@@ -203,9 +171,6 @@ public:
|
||||
stream.Write((const char*)&DataSize, sizeof(DataSize));
|
||||
// size_t ListSize;
|
||||
stream.Write((const char*)&ListSize, sizeof(ListSize));
|
||||
// uint64_t Flags;
|
||||
uint64_t WrittenFlags = FLAG_Shared; // on disk format always has the Shared flag
|
||||
stream.Write((const char*)&WrittenFlags, sizeof(WrittenFlags));
|
||||
|
||||
// inline data
|
||||
stream.Write((const char*)GetData(), DataSize);
|
||||
@@ -226,10 +191,6 @@ public:
|
||||
// size_t ListSize;
|
||||
memcpy(ptr, &ListSize, sizeof(ListSize));
|
||||
ptr += sizeof(ListSize);
|
||||
// uint64_t Flags;
|
||||
uint64_t WrittenFlags = FLAG_Shared; // on disk format always has the Shared flag
|
||||
memcpy(ptr, &WrittenFlags, sizeof(WrittenFlags));
|
||||
ptr += sizeof(WrittenFlags);
|
||||
|
||||
// inline data
|
||||
memcpy(ptr, (const void*)GetData(), DataSize);
|
||||
@@ -240,15 +201,10 @@ public:
|
||||
|
||||
[[nodiscard]]
|
||||
size_t GetInlineSize() const {
|
||||
static_assert(sizeof(*this) == 40);
|
||||
static_assert(sizeof(*this) == 32);
|
||||
return sizeof(*this) + DataSize + ListSize;
|
||||
}
|
||||
|
||||
[[nodiscard]]
|
||||
IRListView* CreateCopy() {
|
||||
return new IRListView(this, true);
|
||||
}
|
||||
|
||||
[[nodiscard]]
|
||||
size_t GetDataSize() const {
|
||||
return DataSize;
|
||||
@@ -263,36 +219,12 @@ public:
|
||||
}
|
||||
|
||||
[[nodiscard]]
|
||||
bool IsCopy() const {
|
||||
return (Flags & FLAG_IsCopy) != 0;
|
||||
}
|
||||
void SetCopy(bool Set) {
|
||||
if (Set) {
|
||||
Flags |= FLAG_IsCopy;
|
||||
} else {
|
||||
Flags &= ~FLAG_IsCopy;
|
||||
}
|
||||
}
|
||||
|
||||
[[nodiscard]]
|
||||
bool IsShared() const {
|
||||
return (Flags & FLAG_Shared) != 0;
|
||||
}
|
||||
void SetShared(bool Set) {
|
||||
if (Set) {
|
||||
Flags |= FLAG_Shared;
|
||||
} else {
|
||||
Flags &= ~FLAG_Shared;
|
||||
}
|
||||
}
|
||||
|
||||
[[nodiscard]]
|
||||
NodeID GetID(const OrderedNode* Node) const {
|
||||
NodeID GetID(const Ref Node) const {
|
||||
return Node->Wrapped(GetListData()).ID();
|
||||
}
|
||||
|
||||
[[nodiscard]]
|
||||
OrderedNode* GetHeaderNode() const {
|
||||
Ref GetHeaderNode() const {
|
||||
OrderedNodeWrapper Wrapped;
|
||||
Wrapped.NodeOffset = sizeof(OrderedNode);
|
||||
return Wrapped.GetNode(GetListData());
|
||||
@@ -305,7 +237,7 @@ public:
|
||||
|
||||
template<typename T>
|
||||
[[nodiscard]]
|
||||
T* GetOp(OrderedNode* Node) const {
|
||||
T* GetOp(Ref Node) const {
|
||||
auto OpHeader = Node->Op(GetData());
|
||||
auto Op = OpHeader->template CW<T>();
|
||||
|
||||
@@ -326,13 +258,13 @@ public:
|
||||
}
|
||||
|
||||
[[nodiscard]]
|
||||
OrderedNode* GetNode(OrderedNodeWrapper Wrapper) const {
|
||||
Ref GetNode(OrderedNodeWrapper Wrapper) const {
|
||||
return Wrapper.GetNode(GetListData());
|
||||
}
|
||||
|
||||
///< Gets an OrderedNode from the IRListView as an OrderedNodeWrapper.
|
||||
[[nodiscard]]
|
||||
OrderedNodeWrapper WrapNode(OrderedNode* Node) const {
|
||||
OrderedNodeWrapper WrapNode(Ref Node) const {
|
||||
return Node->Wrapped(GetListData());
|
||||
}
|
||||
|
||||
@@ -405,7 +337,7 @@ public:
|
||||
}
|
||||
|
||||
[[nodiscard]]
|
||||
CodeRange GetCode(const OrderedNode* block) const {
|
||||
CodeRange GetCode(const Ref block) const {
|
||||
return CodeRange(this, block->Wrapped(GetListData()));
|
||||
}
|
||||
|
||||
@@ -450,7 +382,7 @@ public:
|
||||
}
|
||||
|
||||
[[nodiscard]]
|
||||
iterator at(const OrderedNode* Node) const noexcept {
|
||||
iterator at(const Ref Node) const noexcept {
|
||||
const auto ListData = GetListData();
|
||||
auto Wrapped = Node->Wrapped(ListData);
|
||||
return iterator(ListData, GetData(), Wrapped);
|
||||
@@ -471,15 +403,17 @@ private:
|
||||
void* ListDataInternal;
|
||||
size_t DataSize;
|
||||
size_t ListSize;
|
||||
uint64_t Flags {0};
|
||||
uint8_t InlineData[0];
|
||||
};
|
||||
|
||||
struct IRListViewDeleter {
|
||||
void operator()(IRListView* r) {
|
||||
if (!r->IsShared()) {
|
||||
delete r;
|
||||
}
|
||||
}
|
||||
class IRStorageBase {
|
||||
public:
|
||||
virtual ~IRStorageBase() = default;
|
||||
|
||||
// Optional RA data. Returns nullptr if none present
|
||||
virtual const RegisterAllocationData* RAData() = 0;
|
||||
|
||||
virtual IRListView GetIRView() = 0;
|
||||
};
|
||||
|
||||
} // namespace FEXCore::IR
|
||||
@@ -66,59 +66,39 @@ void PassManager::Finalize() {
|
||||
}
|
||||
}
|
||||
|
||||
void PassManager::AddDefaultPasses(FEXCore::Context::ContextImpl* ctx, bool InlineConstants) {
|
||||
void PassManager::AddDefaultPasses(FEXCore::Context::ContextImpl* ctx) {
|
||||
FEX_CONFIG_OPT(DisablePasses, O0);
|
||||
|
||||
if (!DisablePasses()) {
|
||||
InsertPass(CreateContextLoadStoreElimination(ctx->HostFeatures.SupportsAVX));
|
||||
|
||||
if (Is64BitMode()) {
|
||||
// This needs to run after RCLSE
|
||||
// This only matters for 64-bit code since these instructions don't exist in 32-bit
|
||||
InsertPass(CreateLongDivideEliminationPass());
|
||||
}
|
||||
|
||||
InsertPass(CreateDeadStoreElimination(ctx->HostFeatures.SupportsAVX));
|
||||
InsertPass(CreatePassDeadCodeElimination());
|
||||
InsertPass(CreateConstProp(InlineConstants, ctx->HostFeatures.SupportsTSOImm9, Is64BitMode()));
|
||||
|
||||
InsertPass(CreateContextLoadStoreElimination(ctx->HostFeatures.SupportsAVX && ctx->HostFeatures.SupportsSVE256));
|
||||
InsertPass(CreateDeadStoreElimination());
|
||||
InsertPass(CreateConstProp(ctx->HostFeatures.SupportsTSOImm9, &ctx->CPUID));
|
||||
InsertPass(CreateDeadFlagCalculationEliminination());
|
||||
|
||||
InsertPass(CreateInlineCallOptimization(&ctx->CPUID));
|
||||
InsertPass(CreatePassDeadCodeElimination());
|
||||
}
|
||||
|
||||
// If the IR is compacted post-RA then the node indexing gets messed up and the backend isn't able to find the register assigned to a node
|
||||
// Compact before IR, don't worry about RA generating spills/fills
|
||||
InsertPass(CreateIRCompaction(ctx->OpDispatcherAllocator), "Compaction");
|
||||
}
|
||||
|
||||
void PassManager::AddDefaultValidationPasses() {
|
||||
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
|
||||
InsertValidationPass(Validation::CreateIRValidation(), "IRValidation");
|
||||
InsertValidationPass(Validation::CreateRAValidation());
|
||||
InsertValidationPass(Validation::CreateValueDominanceValidation());
|
||||
#endif
|
||||
}
|
||||
|
||||
void PassManager::InsertRegisterAllocationPass(bool SupportsAVX) {
|
||||
InsertPass(IR::CreateRegisterAllocationPass(GetPass("Compaction"), SupportsAVX), "RA");
|
||||
void PassManager::InsertRegisterAllocationPass() {
|
||||
InsertPass(IR::CreateRegisterAllocationPass(), "RA");
|
||||
}
|
||||
|
||||
bool PassManager::Run(IREmitter* IREmit) {
|
||||
void PassManager::Run(IREmitter* IREmit) {
|
||||
FEXCORE_PROFILE_SCOPED("PassManager::Run");
|
||||
|
||||
bool Changed = false;
|
||||
for (const auto& Pass : Passes) {
|
||||
Changed |= Pass->Run(IREmit);
|
||||
Pass->Run(IREmit);
|
||||
}
|
||||
|
||||
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
|
||||
for (const auto& Pass : ValidationPasses) {
|
||||
Changed |= Pass->Run(IREmit);
|
||||
Pass->Run(IREmit);
|
||||
}
|
||||
#endif
|
||||
|
||||
return Changed;
|
||||
}
|
||||
} // namespace FEXCore::IR
|
||||
@@ -32,7 +32,7 @@ class IREmitter;
|
||||
class Pass {
|
||||
public:
|
||||
virtual ~Pass() = default;
|
||||
virtual bool Run(IREmitter* IREmit) = 0;
|
||||
virtual void Run(IREmitter* IREmit) = 0;
|
||||
|
||||
void RegisterPassManager(PassManager* _Manager) {
|
||||
Manager = _Manager;
|
||||
@@ -43,9 +43,9 @@ protected:
|
||||
};
|
||||
|
||||
class PassManager final {
|
||||
friend class InlineCallOptimization;
|
||||
friend class ConstProp;
|
||||
public:
|
||||
void AddDefaultPasses(FEXCore::Context::ContextImpl* ctx, bool InlineConstants);
|
||||
void AddDefaultPasses(FEXCore::Context::ContextImpl* ctx);
|
||||
void AddDefaultValidationPasses();
|
||||
Pass* InsertPass(fextl::unique_ptr<Pass> Pass, fextl::string Name = "") {
|
||||
auto PassPtr = InsertAt(Passes.end(), std::move(Pass))->get();
|
||||
@@ -56,9 +56,9 @@ public:
|
||||
return PassPtr;
|
||||
}
|
||||
|
||||
void InsertRegisterAllocationPass(bool SupportsAVX);
|
||||
void InsertRegisterAllocationPass();
|
||||
|
||||
bool Run(IREmitter* IREmit);
|
||||
void Run(IREmitter* IREmit);
|
||||
|
||||
bool HasPass(fextl::string Name) const {
|
||||
return NameToPassMaping.contains(Name);
|
||||
|
||||
@@ -16,20 +16,15 @@ class Pass;
|
||||
class RegisterAllocationPass;
|
||||
class RegisterAllocationData;
|
||||
|
||||
fextl::unique_ptr<FEXCore::IR::Pass> CreateConstProp(bool InlineConstants, bool SupportsTSOImm9, bool Is64BitMode);
|
||||
fextl::unique_ptr<FEXCore::IR::Pass> CreateContextLoadStoreElimination(bool SupportsAVX);
|
||||
fextl::unique_ptr<FEXCore::IR::Pass> CreateInlineCallOptimization(const FEXCore::CPUIDEmu* CPUID);
|
||||
fextl::unique_ptr<FEXCore::IR::Pass> CreateConstProp(bool SupportsTSOImm9, const FEXCore::CPUIDEmu* CPUID);
|
||||
fextl::unique_ptr<FEXCore::IR::Pass> CreateContextLoadStoreElimination(bool SupportsSVE256);
|
||||
fextl::unique_ptr<FEXCore::IR::Pass> CreateDeadFlagCalculationEliminination();
|
||||
fextl::unique_ptr<FEXCore::IR::Pass> CreateDeadStoreElimination(bool SupportsAVX);
|
||||
fextl::unique_ptr<FEXCore::IR::Pass> CreatePassDeadCodeElimination();
|
||||
fextl::unique_ptr<FEXCore::IR::Pass> CreateIRCompaction(FEXCore::Utils::IntrusivePooledAllocator& Allocator);
|
||||
fextl::unique_ptr<FEXCore::IR::RegisterAllocationPass> CreateRegisterAllocationPass(FEXCore::IR::Pass* CompactionPass, bool SupportsAVX);
|
||||
fextl::unique_ptr<FEXCore::IR::Pass> CreateLongDivideEliminationPass();
|
||||
fextl::unique_ptr<FEXCore::IR::Pass> CreateDeadStoreElimination();
|
||||
fextl::unique_ptr<FEXCore::IR::RegisterAllocationPass> CreateRegisterAllocationPass();
|
||||
|
||||
namespace Validation {
|
||||
fextl::unique_ptr<FEXCore::IR::Pass> CreateIRValidation();
|
||||
fextl::unique_ptr<FEXCore::IR::Pass> CreateRAValidation();
|
||||
fextl::unique_ptr<FEXCore::IR::Pass> CreateValueDominanceValidation();
|
||||
} // namespace Validation
|
||||
|
||||
namespace Debug {
|
||||
|
||||
File diff suppressed because it is too large.
Load diff
@@ -1,113 +0,0 @@
|
||||
// SPDX-License-Identifier: MIT
|
||||
/*
|
||||
$info$
|
||||
tags: ir|opts
|
||||
$end_info$
|
||||
*/
|
||||
|
||||
#include "Interface/IR/IREmitter.h"
|
||||
#include "Interface/IR/PassManager.h"
|
||||
|
||||
#include <FEXCore/IR/IR.h>
|
||||
#include <FEXCore/Utils/Profiler.h>
|
||||
|
||||
#include <memory>
|
||||
|
||||
namespace FEXCore::IR {
|
||||
|
||||
class DeadCodeElimination final : public FEXCore::IR::Pass {
|
||||
bool Run(IREmitter* IREmit) override;
|
||||
|
||||
private:
|
||||
void markUsed(OrderedNodeWrapper* CodeOp, IROp_Header* IROp);
|
||||
};
|
||||
|
||||
bool DeadCodeElimination::Run(IREmitter* IREmit) {
|
||||
FEXCORE_PROFILE_SCOPED("PassManager::DCE");
|
||||
auto CurrentIR = IREmit->ViewIR();
|
||||
bool Changed = false;
|
||||
|
||||
for (auto [BlockNode, BlockHeader] : CurrentIR.GetBlocks()) {
|
||||
|
||||
// Reverse iteration is not yet working with the iterators
|
||||
auto BlockIROp = BlockHeader->CW<FEXCore::IR::IROp_CodeBlock>();
|
||||
|
||||
// We grab these nodes this way so we can iterate easily
|
||||
auto CodeBegin = CurrentIR.at(BlockIROp->Begin);
|
||||
auto CodeLast = CurrentIR.at(BlockIROp->Last);
|
||||
|
||||
while (1) {
|
||||
auto [CodeNode, IROp] = CodeLast();
|
||||
|
||||
bool HasSideEffects = IR::HasSideEffects(IROp->Op);
|
||||
|
||||
switch (IROp->Op) {
|
||||
case OP_SYSCALL:
|
||||
case OP_INLINESYSCALL: {
|
||||
FEXCore::IR::SyscallFlags Flags {};
|
||||
if (IROp->Op == OP_SYSCALL) {
|
||||
auto Op = IROp->C<IR::IROp_Syscall>();
|
||||
Flags = Op->Flags;
|
||||
} else {
|
||||
auto Op = IROp->C<IR::IROp_InlineSyscall>();
|
||||
Flags = Op->Flags;
|
||||
}
|
||||
|
||||
if ((Flags & FEXCore::IR::SyscallFlags::NOSIDEEFFECTS) == FEXCore::IR::SyscallFlags::NOSIDEEFFECTS) {
|
||||
HasSideEffects = false;
|
||||
}
|
||||
|
||||
break;
|
||||
}
|
||||
case OP_ATOMICFETCHADD:
|
||||
case OP_ATOMICFETCHSUB:
|
||||
case OP_ATOMICFETCHAND:
|
||||
case OP_ATOMICFETCHCLR:
|
||||
case OP_ATOMICFETCHOR:
|
||||
case OP_ATOMICFETCHXOR:
|
||||
case OP_ATOMICFETCHNEG: {
|
||||
// If the result of the atomic fetch is completely unused, convert it to a non-fetching atomic operation.
|
||||
if (CodeNode->GetUses() == 0) {
|
||||
switch (IROp->Op) {
|
||||
case OP_ATOMICFETCHADD: IROp->Op = OP_ATOMICADD; break;
|
||||
case OP_ATOMICFETCHSUB: IROp->Op = OP_ATOMICSUB; break;
|
||||
case OP_ATOMICFETCHAND: IROp->Op = OP_ATOMICAND; break;
|
||||
case OP_ATOMICFETCHCLR: IROp->Op = OP_ATOMICCLR; break;
|
||||
case OP_ATOMICFETCHOR: IROp->Op = OP_ATOMICOR; break;
|
||||
case OP_ATOMICFETCHXOR: IROp->Op = OP_ATOMICXOR; break;
|
||||
case OP_ATOMICFETCHNEG: IROp->Op = OP_ATOMICNEG; break;
|
||||
default: FEX_UNREACHABLE;
|
||||
}
|
||||
Changed = true;
|
||||
}
|
||||
break;
|
||||
}
|
||||
default: break;
|
||||
}
|
||||
|
||||
// Skip over anything that has side effects
|
||||
// Use count tracking can't safely remove anything with side effects
|
||||
if (!HasSideEffects) {
|
||||
if (CodeNode->GetUses() == 0) {
|
||||
Changed = true;
|
||||
IREmit->Remove(CodeNode);
|
||||
}
|
||||
}
|
||||
|
||||
if (CodeLast == CodeBegin) {
|
||||
break;
|
||||
}
|
||||
--CodeLast;
|
||||
}
|
||||
}
|
||||
|
||||
return Changed;
|
||||
}
|
||||
|
||||
void DeadCodeElimination::markUsed(OrderedNodeWrapper* CodeOp, IROp_Header* IROp) {}
|
||||
|
||||
fextl::unique_ptr<FEXCore::IR::Pass> CreatePassDeadCodeElimination() {
|
||||
return fextl::make_unique<DeadCodeElimination>();
|
||||
}
|
||||
|
||||
} // namespace FEXCore::IR
|
||||
@@ -75,9 +75,9 @@ struct ContextMemberInfo {
|
||||
uint32_t AccessOffset;
|
||||
uint8_t AccessSize;
|
||||
///< The last value that was loaded or stored.
|
||||
FEXCore::IR::OrderedNode* ValueNode;
|
||||
FEXCore::IR::Ref ValueNode;
|
||||
///< With a store access, the store node that is doing the operation.
|
||||
FEXCore::IR::OrderedNode* StoreNode;
|
||||
FEXCore::IR::Ref StoreNode;
|
||||
};
|
||||
|
||||
struct ContextInfo {
|
||||
@@ -85,9 +85,39 @@ struct ContextInfo {
|
||||
fextl::vector<ContextMemberInfo> ClassificationInfo;
|
||||
};
|
||||
|
||||
static void ClassifyContextStruct(ContextInfo* ContextClassificationInfo, bool SupportsAVX) {
|
||||
static void ClassifyContextStruct(ContextInfo* ContextClassificationInfo, bool SupportsAVX256) {
|
||||
auto ContextClassification = &ContextClassificationInfo->ClassificationInfo;
|
||||
|
||||
ContextClassification->emplace_back(ContextMemberInfo {
|
||||
ContextMemberClassification {
|
||||
offsetof(FEXCore::Core::CPUState, InlineJITBlockHeader),
|
||||
sizeof(FEXCore::Core::CPUState::InlineJITBlockHeader),
|
||||
},
|
||||
LastAccessType::INVALID,
|
||||
FEXCore::IR::InvalidClass,
|
||||
});
|
||||
|
||||
// DeferredSignalRefCount
|
||||
ContextClassification->emplace_back(ContextMemberInfo {
|
||||
ContextMemberClassification {
|
||||
offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount),
|
||||
sizeof(FEXCore::Core::CPUState::DeferredSignalRefCount),
|
||||
},
|
||||
LastAccessType::NONE,
|
||||
FEXCore::IR::InvalidClass,
|
||||
});
|
||||
|
||||
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_XMMS; ++i) {
|
||||
ContextClassification->emplace_back(ContextMemberInfo {
|
||||
ContextMemberClassification {
|
||||
offsetof(FEXCore::Core::CPUState, avx_high[0][0]) + FEXCore::Core::CPUState::XMM_SSE_REG_SIZE * i,
|
||||
FEXCore::Core::CPUState::XMM_SSE_REG_SIZE,
|
||||
},
|
||||
LastAccessType::NONE,
|
||||
FEXCore::IR::InvalidClass,
|
||||
});
|
||||
}
|
||||
|
||||
ContextClassification->emplace_back(ContextMemberInfo {
|
||||
ContextMemberClassification {
|
||||
offsetof(FEXCore::Core::CPUState, rip),
|
||||
@@ -108,6 +138,50 @@ static void ClassifyContextStruct(ContextInfo* ContextClassificationInfo, bool S
|
||||
});
|
||||
}
|
||||
|
||||
ContextClassification->emplace_back(ContextMemberInfo {
|
||||
ContextMemberClassification {
|
||||
offsetof(FEXCore::Core::CPUState, _pad),
|
||||
sizeof(FEXCore::Core::CPUState::_pad),
|
||||
},
|
||||
LastAccessType::INVALID,
|
||||
FEXCore::IR::InvalidClass,
|
||||
});
|
||||
|
||||
static_assert(offsetof(FEXCore::Core::CPUState, xmm.avx.data[0][0]) == 416, "What");
|
||||
static_assert(FEXCore::Core::CPUState::XMM_AVX_REG_SIZE == 32, "What");
|
||||
if (SupportsAVX256) {
|
||||
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_XMMS; ++i) {
|
||||
ContextClassification->emplace_back(ContextMemberInfo {
|
||||
ContextMemberClassification {
|
||||
offsetof(FEXCore::Core::CPUState, xmm.avx.data[0][0]) + FEXCore::Core::CPUState::XMM_AVX_REG_SIZE * i,
|
||||
FEXCore::Core::CPUState::XMM_AVX_REG_SIZE,
|
||||
},
|
||||
LastAccessType::NONE,
|
||||
FEXCore::IR::InvalidClass,
|
||||
});
|
||||
}
|
||||
} else {
|
||||
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_XMMS; ++i) {
|
||||
ContextClassification->emplace_back(ContextMemberInfo {
|
||||
ContextMemberClassification {
|
||||
offsetof(FEXCore::Core::CPUState, xmm.sse.data[0][0]) + FEXCore::Core::CPUState::XMM_SSE_REG_SIZE * i,
|
||||
FEXCore::Core::CPUState::XMM_SSE_REG_SIZE,
|
||||
},
|
||||
LastAccessType::NONE,
|
||||
FEXCore::IR::InvalidClass,
|
||||
});
|
||||
}
|
||||
|
||||
ContextClassification->emplace_back(ContextMemberInfo {
|
||||
ContextMemberClassification {
|
||||
offsetof(FEXCore::Core::CPUState, xmm.sse.pad[0][0]),
|
||||
static_cast<uint16_t>(FEXCore::Core::CPUState::XMM_SSE_REG_SIZE * FEXCore::Core::CPUState::NUM_XMMS),
|
||||
},
|
||||
LastAccessType::INVALID,
|
||||
FEXCore::IR::InvalidClass,
|
||||
});
|
||||
}
|
||||
|
||||
ContextClassification->emplace_back(ContextMemberInfo {
|
||||
ContextMemberClassification {
|
||||
offsetof(FEXCore::Core::CPUState, es_idx),
|
||||
@@ -164,8 +238,8 @@ static void ClassifyContextStruct(ContextInfo* ContextClassificationInfo, bool S
|
||||
|
||||
ContextClassification->emplace_back(ContextMemberInfo {
|
||||
ContextMemberClassification {
|
||||
offsetof(FEXCore::Core::CPUState, _pad),
|
||||
sizeof(FEXCore::Core::CPUState::_pad),
|
||||
offsetof(FEXCore::Core::CPUState, _pad2),
|
||||
sizeof(FEXCore::Core::CPUState::_pad2),
|
||||
},
|
||||
LastAccessType::INVALID,
|
||||
FEXCore::IR::InvalidClass,
|
||||
@@ -225,48 +299,6 @@ static void ClassifyContextStruct(ContextInfo* ContextClassificationInfo, bool S
|
||||
FEXCore::IR::InvalidClass,
|
||||
});
|
||||
|
||||
ContextClassification->emplace_back(ContextMemberInfo {
|
||||
ContextMemberClassification {
|
||||
offsetof(FEXCore::Core::CPUState, InlineJITBlockHeader),
|
||||
sizeof(FEXCore::Core::CPUState::InlineJITBlockHeader),
|
||||
},
|
||||
LastAccessType::INVALID,
|
||||
FEXCore::IR::InvalidClass,
|
||||
});
|
||||
|
||||
if (SupportsAVX) {
|
||||
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_XMMS; ++i) {
|
||||
ContextClassification->emplace_back(ContextMemberInfo {
|
||||
ContextMemberClassification {
|
||||
offsetof(FEXCore::Core::CPUState, xmm.avx.data[0][0]) + FEXCore::Core::CPUState::XMM_AVX_REG_SIZE * i,
|
||||
FEXCore::Core::CPUState::XMM_AVX_REG_SIZE,
|
||||
},
|
||||
LastAccessType::NONE,
|
||||
FEXCore::IR::InvalidClass,
|
||||
});
|
||||
}
|
||||
} else {
|
||||
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_XMMS; ++i) {
|
||||
ContextClassification->emplace_back(ContextMemberInfo {
|
||||
ContextMemberClassification {
|
||||
offsetof(FEXCore::Core::CPUState, xmm.sse.data[0][0]) + FEXCore::Core::CPUState::XMM_SSE_REG_SIZE * i,
|
||||
FEXCore::Core::CPUState::XMM_SSE_REG_SIZE,
|
||||
},
|
||||
LastAccessType::NONE,
|
||||
FEXCore::IR::InvalidClass,
|
||||
});
|
||||
}
|
||||
|
||||
ContextClassification->emplace_back(ContextMemberInfo {
|
||||
ContextMemberClassification {
|
||||
offsetof(FEXCore::Core::CPUState, xmm.sse.pad[0][0]),
|
||||
static_cast<uint16_t>(FEXCore::Core::CPUState::XMM_SSE_REG_SIZE * FEXCore::Core::CPUState::NUM_XMMS),
|
||||
},
|
||||
LastAccessType::INVALID,
|
||||
FEXCore::IR::InvalidClass,
|
||||
});
|
||||
}
|
||||
|
||||
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_FLAGS; ++i) {
|
||||
ContextClassification->emplace_back(ContextMemberInfo {
|
||||
ContextMemberClassification {
|
||||
@@ -337,37 +369,16 @@ static void ClassifyContextStruct(ContextInfo* ContextClassificationInfo, bool S
|
||||
FEXCore::IR::InvalidClass,
|
||||
});
|
||||
|
||||
// _pad2
|
||||
// _pad3
|
||||
ContextClassification->emplace_back(ContextMemberInfo {
|
||||
ContextMemberClassification {
|
||||
offsetof(FEXCore::Core::CPUState, _pad2),
|
||||
sizeof(FEXCore::Core::CPUState::_pad2),
|
||||
offsetof(FEXCore::Core::CPUState, _pad3),
|
||||
sizeof(FEXCore::Core::CPUState::_pad3),
|
||||
},
|
||||
LastAccessType::NONE,
|
||||
FEXCore::IR::InvalidClass,
|
||||
});
|
||||
|
||||
// DeferredSignalRefCount
|
||||
ContextClassification->emplace_back(ContextMemberInfo {
|
||||
ContextMemberClassification {
|
||||
offsetof(FEXCore::Core::CPUState, DeferredSignalRefCount),
|
||||
sizeof(FEXCore::Core::CPUState::DeferredSignalRefCount),
|
||||
},
|
||||
LastAccessType::NONE,
|
||||
FEXCore::IR::InvalidClass,
|
||||
});
|
||||
|
||||
// DeferredSignalFaultAddress
|
||||
ContextClassification->emplace_back(ContextMemberInfo {
|
||||
ContextMemberClassification {
|
||||
offsetof(FEXCore::Core::CPUState, DeferredSignalFaultAddress),
|
||||
sizeof(FEXCore::Core::CPUState::DeferredSignalFaultAddress),
|
||||
},
|
||||
LastAccessType::NONE,
|
||||
FEXCore::IR::InvalidClass,
|
||||
});
|
||||
|
||||
|
||||
[[maybe_unused]] size_t ClassifiedStructSize {};
|
||||
ContextClassificationInfo->Lookup.reserve(sizeof(FEXCore::Core::CPUState));
|
||||
for (auto& it : *ContextClassification) {
|
||||
@@ -387,7 +398,7 @@ static void ClassifyContextStruct(ContextInfo* ContextClassificationInfo, bool S
|
||||
ContextClassificationInfo->Lookup.size(), sizeof(FEXCore::Core::CPUState));
|
||||
}
|
||||
|
||||
static void ResetClassificationAccesses(ContextInfo* ContextClassificationInfo, bool SupportsAVX) {
|
||||
static void ResetClassificationAccesses(ContextInfo* ContextClassificationInfo, bool SupportsAVX256) {
|
||||
auto ContextClassification = &ContextClassificationInfo->ClassificationInfo;
|
||||
|
||||
auto SetAccess = [&](size_t Offset, LastAccessType Access) {
|
||||
@@ -396,12 +407,41 @@ static void ResetClassificationAccesses(ContextInfo* ContextClassificationInfo,
|
||||
ContextClassification->at(Offset).AccessOffset = 0;
|
||||
ContextClassification->at(Offset).StoreNode = nullptr;
|
||||
};
|
||||
|
||||
size_t Offset = 0;
|
||||
|
||||
///< InlineJITBlockHeader
|
||||
SetAccess(Offset++, LastAccessType::INVALID);
|
||||
|
||||
// DeferredSignalRefCount
|
||||
SetAccess(Offset++, LastAccessType::INVALID);
|
||||
|
||||
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_XMMS; ++i) {
|
||||
///< avx_high
|
||||
SetAccess(Offset++, LastAccessType::NONE);
|
||||
}
|
||||
|
||||
// rip
|
||||
SetAccess(Offset++, LastAccessType::NONE);
|
||||
|
||||
///< gregs
|
||||
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_GPRS; ++i) {
|
||||
SetAccess(Offset++, LastAccessType::NONE);
|
||||
}
|
||||
|
||||
// pad
|
||||
SetAccess(Offset++, LastAccessType::NONE);
|
||||
|
||||
// xmm
|
||||
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_XMMS; ++i) {
|
||||
SetAccess(Offset++, LastAccessType::NONE);
|
||||
}
|
||||
|
||||
// xmm_pad
|
||||
if (!SupportsAVX256) {
|
||||
SetAccess(Offset++, LastAccessType::NONE);
|
||||
}
|
||||
|
||||
// Segment indexes
|
||||
SetAccess(Offset++, LastAccessType::NONE);
|
||||
SetAccess(Offset++, LastAccessType::NONE);
|
||||
@@ -410,7 +450,7 @@ static void ResetClassificationAccesses(ContextInfo* ContextClassificationInfo,
|
||||
SetAccess(Offset++, LastAccessType::NONE);
|
||||
SetAccess(Offset++, LastAccessType::NONE);
|
||||
|
||||
// Pad
|
||||
// Pad2
|
||||
SetAccess(Offset++, LastAccessType::INVALID);
|
||||
|
||||
// Segments
|
||||
@@ -421,80 +461,85 @@ static void ResetClassificationAccesses(ContextInfo* ContextClassificationInfo,
|
||||
SetAccess(Offset++, LastAccessType::NONE);
|
||||
SetAccess(Offset++, LastAccessType::NONE);
|
||||
|
||||
// Pad2
|
||||
SetAccess(Offset++, LastAccessType::INVALID);
|
||||
|
||||
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_XMMS; ++i) {
|
||||
SetAccess(Offset++, LastAccessType::NONE);
|
||||
}
|
||||
|
||||
if (!SupportsAVX) {
|
||||
SetAccess(Offset++, LastAccessType::NONE);
|
||||
}
|
||||
|
||||
///< flags
|
||||
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_FLAGS; ++i) {
|
||||
SetAccess(Offset++, LastAccessType::NONE);
|
||||
}
|
||||
|
||||
// PF/AF
|
||||
SetAccess(Offset++, LastAccessType::NONE);
|
||||
///< pf_raw
|
||||
SetAccess(Offset++, LastAccessType::NONE);
|
||||
|
||||
///< af_raw
|
||||
SetAccess(Offset++, LastAccessType::NONE);
|
||||
|
||||
///< mm
|
||||
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_MMS; ++i) {
|
||||
SetAccess(Offset++, LastAccessType::NONE);
|
||||
}
|
||||
|
||||
///< gdt
|
||||
for (size_t i = 0; i < FEXCore::Core::CPUState::NUM_GDTS; ++i) {
|
||||
SetAccess(Offset++, LastAccessType::NONE);
|
||||
}
|
||||
|
||||
SetAccess(Offset++, LastAccessType::NONE);
|
||||
///< FCW
|
||||
SetAccess(Offset++, LastAccessType::NONE);
|
||||
|
||||
SetAccess(Offset++, LastAccessType::INVALID);
|
||||
SetAccess(Offset++, LastAccessType::INVALID);
|
||||
///< AbridgedFTW
|
||||
SetAccess(Offset++, LastAccessType::NONE);
|
||||
|
||||
// pad3
|
||||
SetAccess(Offset++, LastAccessType::INVALID);
|
||||
}
|
||||
|
||||
struct BlockInfo {
|
||||
fextl::vector<FEXCore::IR::OrderedNode*> Predecessors;
|
||||
fextl::vector<FEXCore::IR::OrderedNode*> Successors;
|
||||
fextl::vector<FEXCore::IR::Ref> Predecessors;
|
||||
fextl::vector<FEXCore::IR::Ref> Successors;
|
||||
ContextInfo IncomingClassifiedStruct;
|
||||
ContextInfo OutgoingClassifiedStruct;
|
||||
};
|
||||
|
||||
class RCLSE final : public FEXCore::IR::Pass {
|
||||
public:
|
||||
explicit RCLSE(bool SupportsAVX_)
|
||||
: SupportsAVX {SupportsAVX_} {
|
||||
ClassifyContextStruct(&ClassifiedStruct, SupportsAVX);
|
||||
DCE = FEXCore::IR::CreatePassDeadCodeElimination();
|
||||
explicit RCLSE(bool SupportsAVX256)
|
||||
: SupportsAVX256 {SupportsAVX256} {
|
||||
ClassifyContextStruct(&ClassifiedStruct, SupportsAVX256);
|
||||
}
|
||||
bool Run(FEXCore::IR::IREmitter* IREmit) override;
|
||||
void Run(FEXCore::IR::IREmitter* IREmit) override;
|
||||
private:
|
||||
fextl::unique_ptr<FEXCore::IR::Pass> DCE;
|
||||
|
||||
ContextInfo ClassifiedStruct;
|
||||
fextl::unordered_map<FEXCore::IR::NodeID, BlockInfo> OffsetToBlockMap;
|
||||
|
||||
bool SupportsAVX;
|
||||
bool SupportsAVX256;
|
||||
|
||||
ContextMemberInfo* FindMemberInfo(ContextInfo* ClassifiedInfo, uint32_t Offset, uint8_t Size);
|
||||
ContextMemberInfo* RecordAccess(ContextMemberInfo* Info, FEXCore::IR::RegisterClassType RegClass, uint32_t Offset, uint8_t Size,
|
||||
LastAccessType AccessType, FEXCore::IR::OrderedNode* Node, FEXCore::IR::OrderedNode* StoreNode = nullptr);
|
||||
LastAccessType AccessType, FEXCore::IR::Ref Node, FEXCore::IR::Ref StoreNode = nullptr);
|
||||
ContextMemberInfo* RecordAccess(ContextInfo* ClassifiedInfo, FEXCore::IR::RegisterClassType RegClass, uint32_t Offset, uint8_t Size,
|
||||
LastAccessType AccessType, FEXCore::IR::OrderedNode* Node, FEXCore::IR::OrderedNode* StoreNode = nullptr);
|
||||
LastAccessType AccessType, FEXCore::IR::Ref Node, FEXCore::IR::Ref StoreNode = nullptr);
|
||||
|
||||
bool HandleLoadFlag(FEXCore::IR::IREmitter* IREmit, ContextInfo* LocalInfo, FEXCore::IR::OrderedNode* CodeNode, unsigned Flag);
|
||||
void HandleLoadFlag(FEXCore::IR::IREmitter* IREmit, ContextInfo* LocalInfo, FEXCore::IR::Ref CodeNode, unsigned Flag);
|
||||
|
||||
// Classify context loads and stores.
|
||||
bool ClassifyContextLoad(FEXCore::IR::IREmitter* IREmit, ContextInfo* LocalInfo, FEXCore::IR::RegisterClassType Class, uint32_t Offset,
|
||||
uint8_t Size, FEXCore::IR::OrderedNode* CodeNode, FEXCore::IR::NodeIterator BlockEnd);
|
||||
bool ClassifyContextStore(FEXCore::IR::IREmitter* IREmit, ContextInfo* LocalInfo, FEXCore::IR::RegisterClassType Class, uint32_t Offset,
|
||||
uint8_t Size, FEXCore::IR::OrderedNode* CodeNode, FEXCore::IR::OrderedNode* ValueNode);
|
||||
void ClassifyContextLoad(FEXCore::IR::IREmitter* IREmit, ContextInfo* LocalInfo, FEXCore::IR::RegisterClassType Class, uint32_t Offset,
|
||||
uint8_t Size, FEXCore::IR::Ref CodeNode, FEXCore::IR::NodeIterator BlockEnd);
|
||||
void ClassifyContextStore(FEXCore::IR::IREmitter* IREmit, ContextInfo* LocalInfo, FEXCore::IR::RegisterClassType Class, uint32_t Offset,
|
||||
uint8_t Size, FEXCore::IR::Ref CodeNode, FEXCore::IR::Ref ValueNode);
|
||||
|
||||
// Block local Passes
|
||||
bool RedundantStoreLoadElimination(FEXCore::IR::IREmitter* IREmit);
|
||||
void RedundantStoreLoadElimination(FEXCore::IR::IREmitter* IREmit);
|
||||
|
||||
unsigned OffsetForReg(FEXCore::IR::RegisterClassType Class, unsigned Reg, unsigned Size) {
|
||||
if (Class == FEXCore::IR::FPRClass) {
|
||||
return Size == 32 ? offsetof(FEXCore::Core::CPUState, xmm.avx.data[Reg][0]) : offsetof(FEXCore::Core::CPUState, xmm.sse.data[Reg][0]);
|
||||
} else if (Reg == FEXCore::Core::CPUState::PF_AS_GREG) {
|
||||
return offsetof(FEXCore::Core::CPUState, pf_raw);
|
||||
} else if (Reg == FEXCore::Core::CPUState::AF_AS_GREG) {
|
||||
return offsetof(FEXCore::Core::CPUState, af_raw);
|
||||
} else {
|
||||
return offsetof(FEXCore::Core::CPUState, gregs[Reg]);
|
||||
}
|
||||
}
|
||||
};
|
||||
|
||||
ContextMemberInfo* RCLSE::FindMemberInfo(ContextInfo* ContextClassificationInfo, uint32_t Offset, uint8_t Size) {
|
||||
@@ -502,7 +547,7 @@ ContextMemberInfo* RCLSE::FindMemberInfo(ContextInfo* ContextClassificationInfo,
|
||||
}
|
||||
|
||||
ContextMemberInfo* RCLSE::RecordAccess(ContextMemberInfo* Info, FEXCore::IR::RegisterClassType RegClass, uint32_t Offset, uint8_t Size,
|
||||
LastAccessType AccessType, FEXCore::IR::OrderedNode* ValueNode, FEXCore::IR::OrderedNode* StoreNode) {
|
||||
LastAccessType AccessType, FEXCore::IR::Ref ValueNode, FEXCore::IR::Ref StoreNode) {
|
||||
LOGMAN_THROW_AA_FMT((Offset + Size) <= (Info->Class.Offset + Info->Class.Size), "Access to context item went over member size");
|
||||
LOGMAN_THROW_AA_FMT(Info->Accessed != LastAccessType::INVALID, "Tried to access invalid member");
|
||||
|
||||
@@ -526,13 +571,13 @@ ContextMemberInfo* RCLSE::RecordAccess(ContextMemberInfo* Info, FEXCore::IR::Reg
|
||||
}
|
||||
|
||||
ContextMemberInfo* RCLSE::RecordAccess(ContextInfo* ClassifiedInfo, FEXCore::IR::RegisterClassType RegClass, uint32_t Offset, uint8_t Size,
|
||||
LastAccessType AccessType, FEXCore::IR::OrderedNode* ValueNode, FEXCore::IR::OrderedNode* StoreNode) {
|
||||
LastAccessType AccessType, FEXCore::IR::Ref ValueNode, FEXCore::IR::Ref StoreNode) {
|
||||
ContextMemberInfo* Info = FindMemberInfo(ClassifiedInfo, Offset, Size);
|
||||
return RecordAccess(Info, RegClass, Offset, Size, AccessType, ValueNode, StoreNode);
|
||||
}
|
||||
|
||||
bool RCLSE::ClassifyContextLoad(FEXCore::IR::IREmitter* IREmit, ContextInfo* LocalInfo, FEXCore::IR::RegisterClassType Class,
|
||||
uint32_t Offset, uint8_t Size, FEXCore::IR::OrderedNode* CodeNode, FEXCore::IR::NodeIterator BlockEnd) {
|
||||
void RCLSE::ClassifyContextLoad(FEXCore::IR::IREmitter* IREmit, ContextInfo* LocalInfo, FEXCore::IR::RegisterClassType Class,
|
||||
uint32_t Offset, uint8_t Size, FEXCore::IR::Ref CodeNode, FEXCore::IR::NodeIterator BlockEnd) {
|
||||
auto Info = FindMemberInfo(LocalInfo, Offset, Size);
|
||||
ContextMemberInfo PreviousMemberInfoCopy = *Info;
|
||||
RecordAccess(Info, Class, Offset, Size, LastAccessType::READ, CodeNode);
|
||||
@@ -544,14 +589,12 @@ bool RCLSE::ClassifyContextLoad(FEXCore::IR::IREmitter* IREmit, ContextInfo* Loc
|
||||
// - Previous access was a store, and we are redundantly loading immediately after the store. Eliminating the store.
|
||||
IREmit->ReplaceAllUsesWithRange(CodeNode, PreviousMemberInfoCopy.ValueNode, IREmit->GetIterator(IREmit->WrapNode(CodeNode)), BlockEnd);
|
||||
RecordAccess(Info, Class, Offset, Size, LastAccessType::READ, PreviousMemberInfoCopy.ValueNode);
|
||||
return true;
|
||||
}
|
||||
// TODO: Optimize the case of partial loads.
|
||||
return false;
|
||||
}
|
||||
|
||||
bool RCLSE::ClassifyContextStore(FEXCore::IR::IREmitter* IREmit, ContextInfo* LocalInfo, FEXCore::IR::RegisterClassType Class,
|
||||
uint32_t Offset, uint8_t Size, FEXCore::IR::OrderedNode* CodeNode, FEXCore::IR::OrderedNode* ValueNode) {
|
||||
void RCLSE::ClassifyContextStore(FEXCore::IR::IREmitter* IREmit, ContextInfo* LocalInfo, FEXCore::IR::RegisterClassType Class,
|
||||
uint32_t Offset, uint8_t Size, FEXCore::IR::Ref CodeNode, FEXCore::IR::Ref ValueNode) {
|
||||
auto Info = FindMemberInfo(LocalInfo, Offset, Size);
|
||||
ContextMemberInfo PreviousMemberInfoCopy = *Info;
|
||||
RecordAccess(Info, Class, Offset, Size, LastAccessType::WRITE, ValueNode, CodeNode);
|
||||
@@ -563,15 +606,13 @@ bool RCLSE::ClassifyContextStore(FEXCore::IR::IREmitter* IREmit, ContextInfo* Lo
|
||||
// Revisit when the new RA lands.
|
||||
#if 0
|
||||
IREmit->Remove(PreviousMemberInfoCopy.StoreNode);
|
||||
return true;
|
||||
#endif
|
||||
}
|
||||
|
||||
// TODO: Optimize the case of partial stores.
|
||||
return false;
|
||||
}
|
||||
|
||||
bool RCLSE::HandleLoadFlag(FEXCore::IR::IREmitter* IREmit, ContextInfo* LocalInfo, FEXCore::IR::OrderedNode* CodeNode, unsigned Flag) {
|
||||
void RCLSE::HandleLoadFlag(FEXCore::IR::IREmitter* IREmit, ContextInfo* LocalInfo, FEXCore::IR::Ref CodeNode, unsigned Flag) {
|
||||
const auto FlagOffset = offsetof(FEXCore::Core::CPUState, flags[Flag]);
|
||||
auto Info = FindMemberInfo(LocalInfo, FlagOffset, 1);
|
||||
LastAccessType LastAccess = Info->Accessed;
|
||||
@@ -582,14 +623,10 @@ bool RCLSE::HandleLoadFlag(FEXCore::IR::IREmitter* IREmit, ContextInfo* LocalInf
|
||||
IREmit->SetWriteCursor(CodeNode);
|
||||
IREmit->ReplaceAllUsesWith(CodeNode, LastValueNode);
|
||||
RecordAccess(Info, FEXCore::IR::GPRClass, FlagOffset, 1, LastAccessType::READ, LastValueNode);
|
||||
return true;
|
||||
} else if (IsReadAccess(LastAccess)) {
|
||||
IREmit->ReplaceAllUsesWith(CodeNode, LastValueNode);
|
||||
RecordAccess(Info, FEXCore::IR::GPRClass, FlagOffset, 1, LastAccessType::READ, LastValueNode);
|
||||
return true;
|
||||
}
|
||||
|
||||
return false;
|
||||
}
|
||||
|
||||
/**
|
||||
@@ -624,11 +661,10 @@ bool RCLSE::HandleLoadFlag(FEXCore::IR::IREmitter* IREmit, ContextInfo* LocalInf
|
||||
* (%%176) StoreContext %175 i128, 0x10, 0xa0
|
||||
|
||||
*/
|
||||
bool RCLSE::RedundantStoreLoadElimination(FEXCore::IR::IREmitter* IREmit) {
|
||||
void RCLSE::RedundantStoreLoadElimination(FEXCore::IR::IREmitter* IREmit) {
|
||||
using namespace FEXCore;
|
||||
using namespace FEXCore::IR;
|
||||
|
||||
bool Changed = false;
|
||||
auto CurrentIR = IREmit->ViewIR();
|
||||
auto OriginalWriteCursor = IREmit->GetWriteCursor();
|
||||
|
||||
@@ -640,21 +676,24 @@ bool RCLSE::RedundantStoreLoadElimination(FEXCore::IR::IREmitter* IREmit) {
|
||||
auto BlockOp = BlockHeader->CW<FEXCore::IR::IROp_CodeBlock>();
|
||||
auto BlockEnd = IREmit->GetIterator(BlockOp->Last);
|
||||
|
||||
ResetClassificationAccesses(&LocalInfo, SupportsAVX);
|
||||
ResetClassificationAccesses(&LocalInfo, SupportsAVX256);
|
||||
|
||||
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
|
||||
if (IROp->Op == OP_STORECONTEXT) {
|
||||
auto Op = IROp->CW<IR::IROp_StoreContext>();
|
||||
Changed |= ClassifyContextStore(IREmit, &LocalInfo, Op->Class, Op->Offset, IROp->Size, CodeNode, CurrentIR.GetNode(Op->Value));
|
||||
ClassifyContextStore(IREmit, &LocalInfo, Op->Class, Op->Offset, IROp->Size, CodeNode, CurrentIR.GetNode(Op->Value));
|
||||
} else if (IROp->Op == OP_STOREREGISTER) {
|
||||
auto Op = IROp->CW<IR::IROp_StoreRegister>();
|
||||
Changed |= ClassifyContextStore(IREmit, &LocalInfo, Op->Class, Op->Offset, IROp->Size, CodeNode, CurrentIR.GetNode(Op->Value));
|
||||
auto Offset = OffsetForReg(Op->Class, Op->Reg, IROp->Size);
|
||||
|
||||
ClassifyContextStore(IREmit, &LocalInfo, Op->Class, Offset, IROp->Size, CodeNode, CurrentIR.GetNode(Op->Value));
|
||||
} else if (IROp->Op == OP_LOADREGISTER) {
|
||||
auto Op = IROp->CW<IR::IROp_LoadRegister>();
|
||||
Changed |= ClassifyContextLoad(IREmit, &LocalInfo, Op->Class, Op->Offset, IROp->Size, CodeNode, BlockEnd);
|
||||
auto Offset = OffsetForReg(Op->Class, Op->Reg, IROp->Size);
|
||||
ClassifyContextLoad(IREmit, &LocalInfo, Op->Class, Offset, IROp->Size, CodeNode, BlockEnd);
|
||||
} else if (IROp->Op == OP_LOADCONTEXT) {
|
||||
auto Op = IROp->CW<IR::IROp_LoadContext>();
|
||||
Changed |= ClassifyContextLoad(IREmit, &LocalInfo, Op->Class, Op->Offset, IROp->Size, CodeNode, BlockEnd);
|
||||
ClassifyContextLoad(IREmit, &LocalInfo, Op->Class, Op->Offset, IROp->Size, CodeNode, BlockEnd);
|
||||
} else if (IROp->Op == OP_STOREFLAG) {
|
||||
const auto Op = IROp->CW<IR::IROp_StoreFlag>();
|
||||
const auto FlagOffset = offsetof(FEXCore::Core::CPUState, flags[0]) + Op->Flag;
|
||||
@@ -665,7 +704,6 @@ bool RCLSE::RedundantStoreLoadElimination(FEXCore::IR::IREmitter* IREmit) {
|
||||
// Flags don't alias, so we can take the simple route here. Kill any flags that have been overwritten
|
||||
if (LastStoreNode != nullptr) {
|
||||
IREmit->Remove(LastStoreNode);
|
||||
Changed = true;
|
||||
}
|
||||
} else if (IROp->Op == OP_INVALIDATEFLAGS) {
|
||||
auto Op = IROp->CW<IR::IROp_InvalidateFlags>();
|
||||
@@ -686,15 +724,14 @@ bool RCLSE::RedundantStoreLoadElimination(FEXCore::IR::IREmitter* IREmit) {
|
||||
RecordAccess(&LocalInfo, FEXCore::IR::GPRClass, FlagOffset, 1, LastAccessType::WRITE, IREmit->_Constant(0), CodeNode);
|
||||
|
||||
IREmit->Remove(LastStoreNode);
|
||||
Changed = true;
|
||||
}
|
||||
}
|
||||
} else if (IROp->Op == OP_LOADFLAG) {
|
||||
const auto Op = IROp->CW<IR::IROp_LoadFlag>();
|
||||
|
||||
Changed |= HandleLoadFlag(IREmit, &LocalInfo, CodeNode, Op->Flag);
|
||||
HandleLoadFlag(IREmit, &LocalInfo, CodeNode, Op->Flag);
|
||||
} else if (IROp->Op == OP_LOADDF) {
|
||||
Changed |= HandleLoadFlag(IREmit, &LocalInfo, CodeNode, X86State::RFLAG_DF_RAW_LOC);
|
||||
HandleLoadFlag(IREmit, &LocalInfo, CodeNode, X86State::RFLAG_DF_RAW_LOC);
|
||||
} else if (IROp->Op == OP_SYSCALL || IROp->Op == OP_INLINESYSCALL) {
|
||||
FEXCore::IR::SyscallFlags Flags {};
|
||||
if (IROp->Op == OP_SYSCALL) {
|
||||
@@ -707,39 +744,29 @@ bool RCLSE::RedundantStoreLoadElimination(FEXCore::IR::IREmitter* IREmit) {
|
||||
|
||||
if ((Flags & FEXCore::IR::SyscallFlags::OPTIMIZETHROUGH) != FEXCore::IR::SyscallFlags::OPTIMIZETHROUGH) {
|
||||
// We can't track through these
|
||||
ResetClassificationAccesses(&LocalInfo, SupportsAVX);
|
||||
ResetClassificationAccesses(&LocalInfo, SupportsAVX256);
|
||||
}
|
||||
} else if (IROp->Op == OP_STORECONTEXTINDEXED || IROp->Op == OP_LOADCONTEXTINDEXED || IROp->Op == OP_BREAK) {
|
||||
// We can't track through these
|
||||
ResetClassificationAccesses(&LocalInfo, SupportsAVX);
|
||||
ResetClassificationAccesses(&LocalInfo, SupportsAVX256);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
IREmit->SetWriteCursor(OriginalWriteCursor);
|
||||
|
||||
return Changed;
|
||||
}
|
||||
|
||||
bool RCLSE::Run(FEXCore::IR::IREmitter* IREmit) {
|
||||
void RCLSE::Run(FEXCore::IR::IREmitter* IREmit) {
|
||||
FEXCORE_PROFILE_SCOPED("PassManager::RCLSE");
|
||||
bool Changed = false;
|
||||
|
||||
// Run up to 5 times
|
||||
for (int i = 0; i < 5 && RedundantStoreLoadElimination(IREmit); i++) {
|
||||
Changed = true;
|
||||
DCE->Run(IREmit);
|
||||
}
|
||||
|
||||
return Changed;
|
||||
RedundantStoreLoadElimination(IREmit);
|
||||
}
|
||||
|
||||
} // namespace
|
||||
|
||||
namespace FEXCore::IR {
|
||||
|
||||
fextl::unique_ptr<FEXCore::IR::Pass> CreateContextLoadStoreElimination(bool SupportsAVX) {
|
||||
return fextl::make_unique<RCLSE>(SupportsAVX);
|
||||
fextl::unique_ptr<FEXCore::IR::Pass> CreateContextLoadStoreElimination(bool SupportsAVX256) {
|
||||
return fextl::make_unique<RCLSE>(SupportsAVX256);
|
||||
}
|
||||
|
||||
} // namespace FEXCore::IR
|
||||
@@ -14,7 +14,6 @@ $end_info$
|
||||
#include <FEXCore/IR/IR.h>
|
||||
#include <FEXCore/Utils/LogManager.h>
|
||||
#include <FEXCore/Utils/Profiler.h>
|
||||
#include <FEXCore/fextl/unordered_map.h>
|
||||
|
||||
#include <memory>
|
||||
#include <stddef.h>
|
||||
@@ -24,221 +23,81 @@ namespace FEXCore::IR {
|
||||
|
||||
constexpr int PropagationRounds = 5;
|
||||
|
||||
// Return a bit representing a single GPR or FPR.
|
||||
static inline uint64_t RegBit(RegisterClassType Class, uint32_t Reg) {
|
||||
uint32_t AdjustedReg = (Class == FPRClass) ? (32 + Reg) : Reg;
|
||||
|
||||
return 1UL << AdjustedReg;
|
||||
}
|
||||
|
||||
class DeadStoreElimination final : public FEXCore::IR::Pass {
|
||||
public:
|
||||
explicit DeadStoreElimination(bool SupportsAVX_)
|
||||
: SupportsAVX {SupportsAVX_} {}
|
||||
|
||||
bool Run(IREmitter* IREmit) override;
|
||||
|
||||
private:
|
||||
bool SupportsAVX;
|
||||
|
||||
bool IsFPR(uint32_t Offset) const {
|
||||
const auto [begin, end] = [this]() -> std::pair<ptrdiff_t, ptrdiff_t> {
|
||||
if (SupportsAVX) {
|
||||
return {offsetof(FEXCore::Core::CpuStateFrame, State.xmm.avx.data[0][0]),
|
||||
offsetof(FEXCore::Core::CpuStateFrame, State.xmm.avx.data[16][0])};
|
||||
} else {
|
||||
return {offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[0][0]),
|
||||
offsetof(FEXCore::Core::CpuStateFrame, State.xmm.sse.data[16][0])};
|
||||
}
|
||||
}();
|
||||
|
||||
if (Offset < begin || Offset >= end) {
|
||||
return false;
|
||||
}
|
||||
|
||||
return true;
|
||||
}
|
||||
|
||||
bool IsTrackedWriteFPR(uint32_t Offset, uint8_t Size) const {
|
||||
if (Size != 16 && Size != 8 && Size != 4) {
|
||||
return false;
|
||||
}
|
||||
if (Offset & 15) {
|
||||
return false;
|
||||
}
|
||||
|
||||
return IsFPR(Offset);
|
||||
}
|
||||
|
||||
uint64_t FPRBit(uint32_t Offset, uint32_t Size) const {
|
||||
if (!IsFPR(Offset)) {
|
||||
return 0;
|
||||
}
|
||||
|
||||
const auto begin = offsetof(Core::CpuStateFrame, State.xmm.avx.data[0][0]);
|
||||
|
||||
const auto regSize = SupportsAVX ? Core::CPUState::XMM_AVX_REG_SIZE : Core::CPUState::XMM_SSE_REG_SIZE;
|
||||
const auto regn = (Offset - begin) / regSize;
|
||||
const auto bitn = regn * 3;
|
||||
|
||||
if (!IsTrackedWriteFPR(Offset, Size)) {
|
||||
return 7UL << (bitn);
|
||||
}
|
||||
|
||||
if (Size == 16) {
|
||||
return 7UL << (bitn);
|
||||
} else if (Size == 8) {
|
||||
return 3UL << (bitn);
|
||||
} else if (Size == 4) {
|
||||
return 1UL << (bitn);
|
||||
} else {
|
||||
LOGMAN_MSG_A_FMT("Unexpected FPR size {}", Size);
|
||||
}
|
||||
|
||||
return 7UL << (bitn); // Return maximum on failure case
|
||||
}
|
||||
void Run(IREmitter* IREmit) override;
|
||||
};
|
||||
|
||||
struct FlagInfo {
|
||||
uint64_t reads {0};
|
||||
uint64_t writes {0};
|
||||
uint64_t kill {0};
|
||||
};
|
||||
|
||||
|
||||
struct GPRInfo {
|
||||
uint32_t reads {0};
|
||||
uint32_t writes {0};
|
||||
uint32_t kill {0};
|
||||
};
|
||||
|
||||
bool IsFullGPR(uint32_t Offset, uint8_t Size) {
|
||||
if (Size != 8) {
|
||||
return false;
|
||||
}
|
||||
if (Offset & 7) {
|
||||
return false;
|
||||
}
|
||||
|
||||
if (Offset < 8 || Offset >= (17 * 8)) {
|
||||
return false;
|
||||
}
|
||||
|
||||
return true;
|
||||
}
|
||||
|
||||
bool IsGPR(uint32_t Offset) {
|
||||
|
||||
if (Offset < 8 || Offset >= (17 * 8)) {
|
||||
return false;
|
||||
}
|
||||
|
||||
return true;
|
||||
}
|
||||
|
||||
uint32_t GPRBit(uint32_t Offset) {
|
||||
if (!IsGPR(Offset)) {
|
||||
return 0;
|
||||
}
|
||||
|
||||
return 1 << ((Offset - 8) / 8);
|
||||
}
|
||||
|
||||
struct FPRInfo {
|
||||
struct ReadWriteKill {
|
||||
uint64_t reads {0};
|
||||
uint64_t writes {0};
|
||||
uint64_t kill {0};
|
||||
};
|
||||
|
||||
struct Info {
|
||||
FlagInfo flag;
|
||||
GPRInfo gpr;
|
||||
FPRInfo fpr;
|
||||
ReadWriteKill flag;
|
||||
ReadWriteKill reg;
|
||||
};
|
||||
|
||||
/**
|
||||
* @brief This is a temporary pass to detect simple multiblock dead flag/gpr/fpr stores
|
||||
* @brief This is a temporary pass to detect simple multiblock dead flag/reg stores
|
||||
*
|
||||
* First pass computes which flags/gprs/fprs are read and written per block
|
||||
* First pass computes which flags/regs are read and written per block
|
||||
*
|
||||
* Second pass computes which flags/gprs/fprs are stored, but overwritten by the next block(s).
|
||||
* It also propagates this information a few times to catch dead flags/gprs/fprs across multiple blocks.
|
||||
* Second pass computes which flags/regs are stored, but overwritten by the next block(s).
|
||||
* It also propagates this information a few times to catch dead flags/regs across multiple blocks.
|
||||
*
|
||||
* Third pass removes the dead stores.
|
||||
*
|
||||
*/
|
||||
bool DeadStoreElimination::Run(IREmitter* IREmit) {
|
||||
void DeadStoreElimination::Run(IREmitter* IREmit) {
|
||||
FEXCORE_PROFILE_SCOPED("PassManager::DSE");
|
||||
|
||||
fextl::unordered_map<OrderedNode*, Info> InfoMap;
|
||||
|
||||
bool Changed = false;
|
||||
auto CurrentIR = IREmit->ViewIR();
|
||||
fextl::vector<Info> InfoMap(CurrentIR.GetSSACount());
|
||||
|
||||
// Pass 1
|
||||
// Compute flags/gprs/fprs read/writes per block
|
||||
// Compute flags/regs read/writes per block
|
||||
// This is conservative and doesn't try to be smart about loads after writes
|
||||
{
|
||||
for (auto [BlockNode, BlockIROp] : CurrentIR.GetBlocks()) {
|
||||
auto& BlockInfo = InfoMap[CurrentIR.GetID(BlockNode).Value];
|
||||
|
||||
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
|
||||
|
||||
auto ClassifyRegisterStore = [this](Info& BlockInfo, uint32_t Offset, uint8_t Size) {
|
||||
//// GPR ////
|
||||
if (IsFullGPR(Offset, Size)) {
|
||||
BlockInfo.gpr.writes |= GPRBit(Offset);
|
||||
} else {
|
||||
BlockInfo.gpr.reads |= GPRBit(Offset);
|
||||
}
|
||||
|
||||
//// FPR ////
|
||||
if (IsTrackedWriteFPR(Offset, Size)) {
|
||||
BlockInfo.fpr.writes |= FPRBit(Offset, Size);
|
||||
} else {
|
||||
BlockInfo.fpr.reads |= FPRBit(Offset, Size);
|
||||
}
|
||||
};
|
||||
|
||||
auto ClassifyRegisterLoad = [this](Info& BlockInfo, uint32_t Offset, uint8_t Size) {
|
||||
//// GPR ////
|
||||
BlockInfo.gpr.reads |= GPRBit(Offset);
|
||||
|
||||
//// FPR ////
|
||||
BlockInfo.fpr.reads |= FPRBit(Offset, Size);
|
||||
};
|
||||
|
||||
//// Flags ////
|
||||
if (IROp->Op == OP_STOREFLAG) {
|
||||
auto Op = IROp->C<IR::IROp_StoreFlag>();
|
||||
|
||||
auto& BlockInfo = InfoMap[BlockNode];
|
||||
|
||||
BlockInfo.flag.writes |= 1UL << Op->Flag;
|
||||
} else if (IROp->Op == OP_INVALIDATEFLAGS) {
|
||||
auto Op = IROp->C<IR::IROp_InvalidateFlags>();
|
||||
|
||||
auto& BlockInfo = InfoMap[BlockNode];
|
||||
|
||||
BlockInfo.flag.writes |= Op->Flags;
|
||||
} else if (IROp->Op == OP_LOADFLAG) {
|
||||
auto Op = IROp->C<IR::IROp_LoadFlag>();
|
||||
|
||||
auto& BlockInfo = InfoMap[BlockNode];
|
||||
|
||||
BlockInfo.flag.reads |= 1UL << Op->Flag;
|
||||
} else if (IROp->Op == OP_LOADDF) {
|
||||
auto& BlockInfo = InfoMap[BlockNode];
|
||||
|
||||
BlockInfo.flag.reads |= 1UL << X86State::RFLAG_DF_RAW_LOC;
|
||||
} else if (IROp->Op == OP_STOREREGISTER) {
|
||||
auto Op = IROp->C<IR::IROp_StoreRegister>();
|
||||
|
||||
auto& BlockInfo = InfoMap[BlockNode];
|
||||
ClassifyRegisterStore(BlockInfo, Op->Offset, IROp->Size);
|
||||
BlockInfo.reg.writes |= RegBit(Op->Class, Op->Reg);
|
||||
} else if (IROp->Op == OP_LOADREGISTER) {
|
||||
auto Op = IROp->C<IR::IROp_LoadRegister>();
|
||||
|
||||
auto& BlockInfo = InfoMap[BlockNode];
|
||||
ClassifyRegisterLoad(BlockInfo, Op->Offset, IROp->Size);
|
||||
BlockInfo.reg.reads |= RegBit(Op->Class, Op->Reg);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Pass 2
|
||||
// Compute flags/gprs/fprs that are stored, but always ovewritten in the next blocks
|
||||
// Compute flags/registers that are stored, but always ovewritten in the next blocks
|
||||
// Propagate the information a few times to eliminate more
|
||||
for (int i = 0; i < PropagationRounds; i++) {
|
||||
for (auto [BlockNode, BlockIROp] : CurrentIR.GetBlocks()) {
|
||||
@@ -248,75 +107,33 @@ bool DeadStoreElimination::Run(IREmitter* IREmit) {
|
||||
|
||||
if (IROp->Op == OP_JUMP) {
|
||||
auto Op = IROp->C<IR::IROp_Jump>();
|
||||
OrderedNode* TargetNode = CurrentIR.GetNode(Op->Header.Args[0]);
|
||||
|
||||
auto& BlockInfo = InfoMap[BlockNode];
|
||||
auto& TargetInfo = InfoMap[TargetNode];
|
||||
|
||||
//// Flags ////
|
||||
auto& BlockInfo = InfoMap[CurrentIR.GetID(BlockNode).Value];
|
||||
auto& TargetInfo = InfoMap[Op->Header.Args[0].ID().Value];
|
||||
|
||||
// stores to remove are written by the next block but not read
|
||||
BlockInfo.flag.kill = TargetInfo.flag.writes & ~(TargetInfo.flag.reads) & ~BlockInfo.flag.reads;
|
||||
BlockInfo.reg.kill = TargetInfo.reg.writes & ~(TargetInfo.reg.reads) & ~BlockInfo.reg.reads;
|
||||
|
||||
// Flags that are written by the next block can be considered as written by this block, if not read
|
||||
BlockInfo.flag.writes |= BlockInfo.flag.kill & ~BlockInfo.flag.reads;
|
||||
|
||||
|
||||
//// GPRs ////
|
||||
|
||||
// stores to remove are written by the next block but not read
|
||||
BlockInfo.gpr.kill = TargetInfo.gpr.writes & ~(TargetInfo.gpr.reads) & ~BlockInfo.gpr.reads;
|
||||
|
||||
// GPRs that are written by the next block can be considered as written by this block, if not read
|
||||
BlockInfo.gpr.writes |= BlockInfo.gpr.kill & ~BlockInfo.gpr.reads;
|
||||
|
||||
|
||||
//// FPRs ////
|
||||
|
||||
// stores to remove are written by the next block but not read
|
||||
BlockInfo.fpr.kill = TargetInfo.fpr.writes & ~(TargetInfo.fpr.reads) & ~BlockInfo.fpr.reads;
|
||||
|
||||
// FPRs that are written by the next block can be considered as written by this block, if not read
|
||||
BlockInfo.fpr.writes |= BlockInfo.fpr.kill & ~BlockInfo.fpr.reads;
|
||||
|
||||
BlockInfo.reg.writes |= BlockInfo.reg.kill & ~BlockInfo.reg.reads;
|
||||
} else if (IROp->Op == OP_CONDJUMP) {
|
||||
auto Op = IROp->C<IR::IROp_CondJump>();
|
||||
|
||||
OrderedNode* TrueTargetNode = CurrentIR.GetNode(Op->TrueBlock);
|
||||
OrderedNode* FalseTargetNode = CurrentIR.GetNode(Op->FalseBlock);
|
||||
|
||||
auto& BlockInfo = InfoMap[BlockNode];
|
||||
auto& TrueTargetInfo = InfoMap[TrueTargetNode];
|
||||
auto& FalseTargetInfo = InfoMap[FalseTargetNode];
|
||||
|
||||
//// Flags ////
|
||||
auto& BlockInfo = InfoMap[CurrentIR.GetID(BlockNode).Value];
|
||||
auto& TrueTargetInfo = InfoMap[Op->TrueBlock.ID().Value];
|
||||
auto& FalseTargetInfo = InfoMap[Op->FalseBlock.ID().Value];
|
||||
|
||||
// stores to remove are written by the next blocks but not read
|
||||
BlockInfo.flag.kill = TrueTargetInfo.flag.writes & ~(TrueTargetInfo.flag.reads) & ~BlockInfo.flag.reads;
|
||||
BlockInfo.reg.kill = TrueTargetInfo.reg.writes & ~(TrueTargetInfo.reg.reads) & ~BlockInfo.reg.reads;
|
||||
|
||||
BlockInfo.flag.kill &= FalseTargetInfo.flag.writes & ~(FalseTargetInfo.flag.reads) & ~BlockInfo.flag.reads;
|
||||
BlockInfo.reg.kill &= FalseTargetInfo.reg.writes & ~(FalseTargetInfo.reg.reads) & ~BlockInfo.reg.reads;
|
||||
|
||||
// Flags that are written by the next blocks can be considered as written by this block, if not read
|
||||
BlockInfo.flag.writes |= BlockInfo.flag.kill & ~BlockInfo.flag.reads;
|
||||
|
||||
|
||||
//// GPRs ////
|
||||
|
||||
// stores to remove are written by the next blocks but not read
|
||||
BlockInfo.gpr.kill = TrueTargetInfo.gpr.writes & ~(TrueTargetInfo.gpr.reads) & ~BlockInfo.gpr.reads;
|
||||
BlockInfo.gpr.kill &= FalseTargetInfo.gpr.writes & ~(FalseTargetInfo.gpr.reads) & ~BlockInfo.gpr.reads;
|
||||
|
||||
// GPRs that are written by the next blocks can be considered as written by this block, if not read
|
||||
BlockInfo.gpr.writes |= BlockInfo.gpr.kill & ~BlockInfo.gpr.reads;
|
||||
|
||||
|
||||
//// FPRs ////
|
||||
|
||||
// stores to remove are written by the next blocks but not read
|
||||
BlockInfo.fpr.kill = TrueTargetInfo.fpr.writes & ~(TrueTargetInfo.fpr.reads) & ~BlockInfo.fpr.reads;
|
||||
BlockInfo.fpr.kill &= FalseTargetInfo.fpr.writes & ~(FalseTargetInfo.fpr.reads) & ~BlockInfo.fpr.reads;
|
||||
|
||||
// FPRs that are written by the next blocks can be considered as written by this block, if not read
|
||||
BlockInfo.fpr.writes |= BlockInfo.fpr.kill & ~BlockInfo.fpr.reads;
|
||||
BlockInfo.reg.writes |= BlockInfo.reg.kill & ~BlockInfo.reg.reads;
|
||||
}
|
||||
}
|
||||
}
|
||||
@@ -325,55 +142,31 @@ bool DeadStoreElimination::Run(IREmitter* IREmit) {
|
||||
// Remove the dead stores
|
||||
{
|
||||
for (auto [BlockNode, BlockIROp] : CurrentIR.GetBlocks()) {
|
||||
auto& BlockInfo = InfoMap[CurrentIR.GetID(BlockNode).Value];
|
||||
|
||||
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
|
||||
|
||||
auto RemoveDeadRegisterStore = [this](FEXCore::IR::IREmitter* IREmit, FEXCore::IR::OrderedNode* CodeNode, Info& BlockInfo,
|
||||
uint32_t Offset, uint8_t Size) -> bool {
|
||||
bool Changed {};
|
||||
//// GPRs ////
|
||||
// If this OP_STOREREGISTER is never read, remove it
|
||||
if (BlockInfo.gpr.kill & GPRBit(Offset)) {
|
||||
IREmit->Remove(CodeNode);
|
||||
Changed = true;
|
||||
}
|
||||
|
||||
//// FPRs ////
|
||||
// If this OP_STOREREGISTER is never read, remove it
|
||||
if ((BlockInfo.fpr.kill & FPRBit(Offset, Size)) == FPRBit(Offset, Size) && (FPRBit(Offset, Size) != 0)) {
|
||||
IREmit->Remove(CodeNode);
|
||||
Changed = true;
|
||||
}
|
||||
|
||||
return Changed;
|
||||
};
|
||||
|
||||
//// Flags ////
|
||||
if (IROp->Op == OP_STOREFLAG) {
|
||||
auto Op = IROp->C<IR::IROp_StoreFlag>();
|
||||
|
||||
auto& BlockInfo = InfoMap[BlockNode];
|
||||
|
||||
// If this StoreFlag is never read, remove it
|
||||
if (BlockInfo.flag.kill & (1UL << Op->Flag)) {
|
||||
IREmit->Remove(CodeNode);
|
||||
Changed = true;
|
||||
}
|
||||
} else if (IROp->Op == OP_STOREREGISTER) {
|
||||
auto Op = IROp->C<IR::IROp_StoreRegister>();
|
||||
|
||||
auto& BlockInfo = InfoMap[BlockNode];
|
||||
|
||||
Changed |= RemoveDeadRegisterStore(IREmit, CodeNode, BlockInfo, Op->Offset, IROp->Size);
|
||||
// If this OP_STOREREGISTER is never read, remove it
|
||||
if (BlockInfo.reg.kill & RegBit(Op->Class, Op->Reg)) {
|
||||
IREmit->Remove(CodeNode);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
return Changed;
|
||||
}
|
||||
|
||||
fextl::unique_ptr<FEXCore::IR::Pass> CreateDeadStoreElimination(bool SupportsAVX) {
|
||||
return fextl::make_unique<DeadStoreElimination>(SupportsAVX);
|
||||
fextl::unique_ptr<FEXCore::IR::Pass> CreateDeadStoreElimination() {
|
||||
return fextl::make_unique<DeadStoreElimination>();
|
||||
}
|
||||
|
||||
} // namespace FEXCore::IR
|
||||
@@ -1,228 +0,0 @@
|
||||
// SPDX-License-Identifier: MIT
|
||||
/*
|
||||
$info$
|
||||
tags: ir|opts
|
||||
desc: Sorts the ssa storage in memory, needed for RA and others
|
||||
$end_info$
|
||||
*/
|
||||
|
||||
#include "Interface/IR/IREmitter.h"
|
||||
#include "Interface/IR/PassManager.h"
|
||||
#include "Interface/Core/OpcodeDispatcher.h"
|
||||
|
||||
#include <FEXCore/IR/IR.h>
|
||||
#include <FEXCore/Utils/LogManager.h>
|
||||
#include <FEXCore/Utils/MathUtils.h>
|
||||
#include <FEXCore/Utils/Profiler.h>
|
||||
#include <FEXCore/fextl/vector.h>
|
||||
|
||||
#include <algorithm>
|
||||
#include <cstdint>
|
||||
#include <cstring>
|
||||
#include <memory>
|
||||
|
||||
namespace FEXCore::IR {
|
||||
|
||||
// struct to avoid zero-initialization
|
||||
struct RemapNode {
|
||||
IR::NodeID NodeID;
|
||||
};
|
||||
|
||||
static_assert(sizeof(RemapNode) == 4);
|
||||
|
||||
class IRCompaction final : public FEXCore::IR::Pass {
|
||||
public:
|
||||
IRCompaction(FEXCore::Utils::IntrusivePooledAllocator& Allocator);
|
||||
bool Run(IREmitter* IREmit) override;
|
||||
|
||||
private:
|
||||
static constexpr size_t AlignSize = 0x2000;
|
||||
OpDispatchBuilder LocalBuilder;
|
||||
fextl::vector<RemapNode> OldToNewRemap;
|
||||
struct CodeBlockData {
|
||||
OrderedNode* OldNode;
|
||||
OrderedNode* NewNode;
|
||||
};
|
||||
|
||||
fextl::vector<CodeBlockData> GeneratedCodeBlocks {};
|
||||
};
|
||||
|
||||
IRCompaction::IRCompaction(FEXCore::Utils::IntrusivePooledAllocator& Allocator)
|
||||
: LocalBuilder {Allocator} {
|
||||
OldToNewRemap.resize(AlignSize);
|
||||
}
|
||||
|
||||
bool IRCompaction::Run(IREmitter* IREmit) {
|
||||
FEXCORE_PROFILE_SCOPED("PassManager::IRCompaction");
|
||||
|
||||
LocalBuilder.ReownOrClaimBuffer();
|
||||
|
||||
auto CurrentIR = IREmit->ViewIR();
|
||||
uint32_t NodeCount = CurrentIR.GetSSACount();
|
||||
|
||||
if (OldToNewRemap.size() < NodeCount) {
|
||||
OldToNewRemap.resize(std::max(OldToNewRemap.size() * 2U, AlignUp(NodeCount, AlignSize)));
|
||||
}
|
||||
#ifndef NDEBUG
|
||||
memset(&OldToNewRemap.at(0), 0xFF, NodeCount * sizeof(RemapNode));
|
||||
#endif
|
||||
|
||||
GeneratedCodeBlocks.clear();
|
||||
|
||||
// Reset our local working list
|
||||
LocalBuilder.ResetWorkingList();
|
||||
auto LocalIR = LocalBuilder.ViewIR();
|
||||
|
||||
uintptr_t LocalListBegin = LocalIR.GetListData();
|
||||
uintptr_t LocalDataBegin = LocalIR.GetData();
|
||||
|
||||
uintptr_t ListBegin = CurrentIR.GetListData();
|
||||
|
||||
auto HeaderNode = CurrentIR.GetHeaderNode();
|
||||
auto HeaderOp = CurrentIR.GetHeader();
|
||||
LOGMAN_THROW_AA_FMT(HeaderOp->Header.Op == OP_IRHEADER, "First op wasn't IRHeader");
|
||||
|
||||
// This compaction pass is something that we need to ensure correct ordering and distances between IROps
|
||||
// Later on we assume that an IROp's SSA value live range is its Node locations
|
||||
//
|
||||
// RA distance calculation is calculated purely on the Node locations
|
||||
// So we need to reorder those
|
||||
//
|
||||
// Additionally there may be some dead ops hanging out in the IR list that are orphaned.
|
||||
// These can also be dropped during this pass
|
||||
|
||||
// First thing is first, we need to do some housekeeping
|
||||
// Create the IRHeader op
|
||||
// Create the codeblocks
|
||||
// Then create all the ops inside the code blocks
|
||||
|
||||
// Zero is always zero(invalid)
|
||||
OldToNewRemap[0].NodeID.Invalidate();
|
||||
auto LocalHeaderOp = LocalBuilder._IRHeader(OrderedNodeWrapper::WrapOffset(0).GetNode(ListBegin), HeaderOp->OriginalRIP,
|
||||
HeaderOp->BlockCount, HeaderOp->NumHostInstructions);
|
||||
|
||||
OldToNewRemap[CurrentIR.GetID(HeaderNode).Value].NodeID = LocalIR.GetID(LocalHeaderOp.Node);
|
||||
|
||||
{
|
||||
// Generate our codeblocks and link them together
|
||||
for (auto [BlockNode, BlockHeader] : CurrentIR.GetBlocks()) {
|
||||
LOGMAN_THROW_AA_FMT(BlockHeader->Op == OP_CODEBLOCK, "IR type failed to be a code block");
|
||||
|
||||
auto LocalBlockIRNode = LocalBuilder._CodeBlock(LocalHeaderOp, LocalHeaderOp); // Use LocalHeaderOp as a dummy arg for now
|
||||
OldToNewRemap[CurrentIR.GetID(BlockNode).Value].NodeID = LocalIR.GetID(LocalBlockIRNode.Node);
|
||||
GeneratedCodeBlocks.emplace_back(CodeBlockData {BlockNode, LocalBlockIRNode});
|
||||
}
|
||||
|
||||
// Link the IRHeader to the first code block
|
||||
LocalHeaderOp.first->Blocks = GeneratedCodeBlocks[0].NewNode->Wrapped(LocalListBegin);
|
||||
}
|
||||
|
||||
{
|
||||
// Copy all of our IR ops over to the new location
|
||||
for (auto& Block : GeneratedCodeBlocks) {
|
||||
|
||||
// Isolate block contents from any previous headers/blocks
|
||||
LocalBuilder.SetWriteCursor(nullptr);
|
||||
|
||||
CodeBlockData FirstNode {};
|
||||
CodeBlockData LastNode {};
|
||||
uint32_t i {};
|
||||
for (auto [CodeNode, IROp] : CurrentIR.GetCode(Block.OldNode)) {
|
||||
const size_t OpSize = FEXCore::IR::GetSize(IROp->Op);
|
||||
const auto CodeID = CurrentIR.GetID(CodeNode);
|
||||
|
||||
// Allocate the ops locally for our local dispatch
|
||||
auto LocalPair = LocalBuilder.AllocateRawOp(OpSize);
|
||||
|
||||
// Copy usage infomation
|
||||
LocalPair.Node->NumUses = CodeNode->GetUses();
|
||||
|
||||
// Copy over the op
|
||||
memcpy(LocalPair.first, IROp, OpSize);
|
||||
|
||||
// Set our map remapper to map the new location
|
||||
// Even nodes that don't have a destination need to be in this map
|
||||
// Need to be able to remap branch targets any other bits
|
||||
OldToNewRemap[CodeID.Value].NodeID = LocalIR.GetID(LocalPair.Node);
|
||||
|
||||
if (i == 0) {
|
||||
FirstNode.OldNode = CodeNode;
|
||||
FirstNode.NewNode = LocalPair.Node;
|
||||
}
|
||||
|
||||
if (IROp->Op == OP_ENDBLOCK) {
|
||||
LastNode.OldNode = CodeNode;
|
||||
LastNode.NewNode = LocalPair.Node;
|
||||
}
|
||||
|
||||
++i;
|
||||
}
|
||||
|
||||
// Set the code block's begin and end correctly
|
||||
auto NewBlockIROp = Block.NewNode->Op(LocalDataBegin)->CW<FEXCore::IR::IROp_CodeBlock>();
|
||||
NewBlockIROp->Begin = FirstNode.NewNode->Wrapped(LocalListBegin);
|
||||
NewBlockIROp->Last = LastNode.NewNode->Wrapped(LocalListBegin);
|
||||
}
|
||||
}
|
||||
|
||||
{
|
||||
// Fixup the arguments of all the IROps
|
||||
for (auto& Block : GeneratedCodeBlocks) {
|
||||
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
|
||||
auto BlockIROp = LocalIR.GetOp<FEXCore::IR::IROp_CodeBlock>(Block.NewNode);
|
||||
LOGMAN_THROW_AA_FMT(BlockIROp->Header.Op == OP_CODEBLOCK, "IR type failed to be a code block");
|
||||
#endif
|
||||
|
||||
for (auto [LocalNode, LocalIROp] : LocalIR.GetCode(Block.NewNode)) {
|
||||
|
||||
// Now that we have the op copied over, we need to modify SSA values to point to the new correct locations
|
||||
// This doesn't use IR::GetRAArgs(Op) because we need to remap all SSA nodes
|
||||
// Including ones that we don't RA
|
||||
const uint8_t NumArgs = IR::GetArgs(LocalIROp->Op);
|
||||
for (uint8_t i = 0; i < NumArgs; ++i) {
|
||||
const auto OldArg = LocalIROp->Args[i].ID();
|
||||
const auto NewArg = OldToNewRemap[OldArg.Value].NodeID;
|
||||
|
||||
#ifndef NDEBUG
|
||||
LOGMAN_THROW_A_FMT(NewArg.Value != UINT32_MAX, "Tried remapping unfound node %{}", OldArg);
|
||||
#endif
|
||||
|
||||
LocalIROp->Args[i].NodeOffset = NewArg.Value * sizeof(OrderedNode);
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// uintptr_t OldListSize = CurrentIR.GetListSize();
|
||||
// uintptr_t OldDataSize = CurrentIR.GetDataSize();
|
||||
|
||||
// uintptr_t NewListSize = LocalIR.GetListSize();
|
||||
// uintptr_t NewDataSize = LocalIR.GetDataSize();
|
||||
|
||||
// if (NewListSize < OldListSize ||
|
||||
// NewDataSize < OldDataSize) {
|
||||
// if (NewListSize < OldListSize) {
|
||||
// LogMan::Msg::DFmt("Shaved {} bytes off the list size", OldListSize - NewListSize);
|
||||
// }
|
||||
// if (NewDataSize < OldDataSize) {
|
||||
// LogMan::Msg::DFmt("Shaved {} bytes off the data size", OldDataSize - NewDataSize);
|
||||
// }
|
||||
// }
|
||||
|
||||
// if (NewListSize > OldListSize ||
|
||||
// NewDataSize > OldDataSize) {
|
||||
// LOGMAN_MSG_A_FMT("Whoa. Compaction made the IR a different size when it shouldn't have. 0x{:x} > 0x{:x} or 0x{:x} > 0x{:x}",
|
||||
// NewListSize, OldListSize, NewDataSize, OldDataSize);
|
||||
// }
|
||||
|
||||
IREmit->CopyData(LocalBuilder);
|
||||
|
||||
LocalBuilder.DelayedDisownBuffer();
|
||||
return true;
|
||||
}
|
||||
|
||||
fextl::unique_ptr<FEXCore::IR::Pass> CreateIRCompaction(FEXCore::Utils::IntrusivePooledAllocator& Allocator) {
|
||||
return fextl::make_unique<IRCompaction>(Allocator);
|
||||
}
|
||||
|
||||
} // namespace FEXCore::IR
|
||||
@@ -18,7 +18,7 @@ namespace FEXCore::IR::Debug {
|
||||
class IRDumper final : public FEXCore::IR::Pass {
|
||||
public:
|
||||
IRDumper();
|
||||
bool Run(IREmitter* IREmit) override;
|
||||
void Run(IREmitter* IREmit) override;
|
||||
|
||||
private:
|
||||
FEX_CONFIG_OPT(DumpIR, DUMPIR);
|
||||
@@ -37,7 +37,7 @@ IRDumper::IRDumper() {
|
||||
}
|
||||
}
|
||||
|
||||
bool IRDumper::Run(IREmitter* IREmit) {
|
||||
void IRDumper::Run(IREmitter* IREmit) {
|
||||
auto RAPass = Manager->GetPass<IR::RegisterAllocationPass>("RA");
|
||||
IR::RegisterAllocationData* RA {};
|
||||
if (RAPass) {
|
||||
@@ -71,8 +71,6 @@ bool IRDumper::Run(IREmitter* IREmit) {
|
||||
LogMan::Msg::IFmt("IR-{} 0x{:x}:\n{}\n@@@@@\n", RA ? "post" : "pre", HeaderOp->OriginalRIP, out.str());
|
||||
}
|
||||
}
|
||||
|
||||
return false;
|
||||
}
|
||||
|
||||
fextl::unique_ptr<FEXCore::IR::Pass> CreateIRDumper() {
|
||||
|
||||
@@ -32,7 +32,7 @@ IRValidation::~IRValidation() {
|
||||
NodeIsLive.Free();
|
||||
}
|
||||
|
||||
bool IRValidation::Run(IREmitter* IREmit) {
|
||||
void IRValidation::Run(IREmitter* IREmit) {
|
||||
FEXCORE_PROFILE_SCOPED("PassManager::IRValidation");
|
||||
|
||||
bool HadError = false;
|
||||
@@ -46,11 +46,12 @@ bool IRValidation::Run(IREmitter* IREmit) {
|
||||
OffsetToBlockMap.clear();
|
||||
EntryBlock = nullptr;
|
||||
|
||||
if (CurrentIR.GetSSACount() > MaxNodes) {
|
||||
NodeIsLive.Realloc(CurrentIR.GetSSACount());
|
||||
uint32_t Count = CurrentIR.GetSSACount();
|
||||
if (Count > MaxNodes) {
|
||||
NodeIsLive.Realloc(Count);
|
||||
}
|
||||
|
||||
fextl::vector<uint32_t> Uses(CurrentIR.GetSSACount(), 0);
|
||||
fextl::vector<uint32_t> Uses(Count, 0);
|
||||
|
||||
#if defined(ASSERTIONS_ENABLED) && ASSERTIONS_ENABLED
|
||||
auto HeaderOp = CurrentIR.GetHeader();
|
||||
@@ -62,8 +63,6 @@ bool IRValidation::Run(IREmitter* IREmit) {
|
||||
RAData = Manager->GetPass<IR::RegisterAllocationPass>("RA")->GetAllocationData();
|
||||
}
|
||||
|
||||
NodeIsLive.Set(1); // IRHEADER
|
||||
|
||||
for (auto [BlockNode, BlockHeader] : CurrentIR.GetBlocks()) {
|
||||
auto BlockIROp = BlockHeader->CW<FEXCore::IR::IROp_CodeBlock>();
|
||||
LOGMAN_THROW_AA_FMT(BlockIROp->Header.Op == OP_CODEBLOCK, "IR type failed to be a code block");
|
||||
@@ -75,6 +74,9 @@ bool IRValidation::Run(IREmitter* IREmit) {
|
||||
const auto BlockID = CurrentIR.GetID(BlockNode);
|
||||
BlockInfo* CurrentBlock = &OffsetToBlockMap.try_emplace(BlockID).first->second;
|
||||
|
||||
// We only allow defs local to a single block, so clear live set per block
|
||||
NodeIsLive.MemClear(Count);
|
||||
|
||||
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
|
||||
const auto ID = CurrentIR.GetID(CodeNode);
|
||||
const uint8_t OpSize = IROp->Size;
|
||||
@@ -125,21 +127,21 @@ bool IRValidation::Run(IREmitter* IREmit) {
|
||||
for (uint32_t i = 0; i < NumArgs; ++i) {
|
||||
OrderedNodeWrapper Arg = IROp->Args[i];
|
||||
const auto ArgID = Arg.ID();
|
||||
|
||||
// Was an argument defined after this node?
|
||||
if (ArgID >= ID) {
|
||||
HadError |= true;
|
||||
Errors << "%" << ID << ": Arg[" << i << "] has definition after use at %" << ArgID << std::endl;
|
||||
}
|
||||
|
||||
if (ArgID.IsValid() && !NodeIsLive.Get(ArgID.Value)) {
|
||||
HadError |= true;
|
||||
Errors << "%" << ID << ": Arg[" << i << "] references dead %" << ArgID << std::endl;
|
||||
}
|
||||
IROps Op = CurrentIR.GetOp<IROp_Header>(Arg)->Op;
|
||||
|
||||
if (ArgID.IsValid()) {
|
||||
Uses[ArgID.Value]++;
|
||||
}
|
||||
|
||||
// We do not validate the location of inline constants because it's
|
||||
// irrelevant, they're ignored by RA and always inlined to where they
|
||||
// need to be. This lets us pool inline constants globally.
|
||||
bool Ignore = (Op == OP_IRHEADER || Op == OP_INLINECONSTANT);
|
||||
|
||||
if (!Ignore && ArgID.IsValid() && !NodeIsLive.Get(ArgID.Value)) {
|
||||
HadError |= true;
|
||||
Errors << "%" << ID << ": Arg[" << i << "] references invalid %" << ArgID << std::endl;
|
||||
}
|
||||
}
|
||||
|
||||
NodeIsLive.Set(ID.Value);
|
||||
@@ -265,8 +267,6 @@ bool IRValidation::Run(IREmitter* IREmit) {
|
||||
Errors.clear();
|
||||
Warnings.clear();
|
||||
}
|
||||
|
||||
return false;
|
||||
}
|
||||
|
||||
fextl::unique_ptr<FEXCore::IR::Pass> CreateIRValidation() {
|
||||
|
||||
@@ -21,7 +21,7 @@ class RAValidation;
|
||||
class IRValidation final : public FEXCore::IR::Pass {
|
||||
public:
|
||||
~IRValidation();
|
||||
bool Run(IREmitter* IREmit) override;
|
||||
void Run(IREmitter* IREmit) override;
|
||||
|
||||
private:
|
||||
|
||||
|
||||
@@ -1,129 +0,0 @@
|
||||
// SPDX-License-Identifier: MIT
|
||||
/*
|
||||
$info$
|
||||
tags: ir|opts
|
||||
desc: Removes unused arguments if known syscall number
|
||||
$end_info$
|
||||
*/
|
||||
|
||||
#include "Interface/Core/CPUID.h"
|
||||
#include "Interface/IR/IREmitter.h"
|
||||
#include "Interface/IR/PassManager.h"
|
||||
|
||||
#include <FEXCore/IR/IR.h>
|
||||
#include <FEXCore/HLE/SyscallHandler.h>
|
||||
#include <FEXCore/Utils/Profiler.h>
|
||||
|
||||
#include <memory>
|
||||
#include <stdint.h>
|
||||
|
||||
namespace FEXCore::IR {
|
||||
|
||||
class InlineCallOptimization final : public FEXCore::IR::Pass {
|
||||
public:
|
||||
InlineCallOptimization(const FEXCore::CPUIDEmu* CPUID)
|
||||
: CPUID {CPUID} {}
|
||||
bool Run(IREmitter* IREmit) override;
|
||||
private:
|
||||
const FEXCore::CPUIDEmu* CPUID;
|
||||
};
|
||||
|
||||
bool InlineCallOptimization::Run(IREmitter* IREmit) {
|
||||
FEXCORE_PROFILE_SCOPED("PassManager::SyscallOpt");
|
||||
|
||||
bool Changed = false;
|
||||
auto CurrentIR = IREmit->ViewIR();
|
||||
|
||||
for (auto [CodeNode, IROp] : CurrentIR.GetAllCode()) {
|
||||
|
||||
if (IROp->Op == FEXCore::IR::OP_SYSCALL) {
|
||||
auto Op = IROp->CW<IR::IROp_Syscall>();
|
||||
|
||||
// Is the first argument a constant?
|
||||
uint64_t Constant;
|
||||
if (IREmit->IsValueConstant(Op->SyscallID, &Constant)) {
|
||||
auto SyscallDef = Manager->SyscallHandler->GetSyscallABI(Constant);
|
||||
auto SyscallFlags = Manager->SyscallHandler->GetSyscallFlags(Constant);
|
||||
|
||||
// Update the syscall flags
|
||||
Op->Flags = SyscallFlags;
|
||||
|
||||
// XXX: Once we have the ability to do real function calls then we can call directly in to the syscall handler
|
||||
if (SyscallDef.NumArgs < FEXCore::HLE::SyscallArguments::MAX_ARGS) {
|
||||
// If the number of args are less than what the IR op supports then we can remove arg usage
|
||||
// We need +1 since we are still passing in syscall number here
|
||||
for (uint8_t Arg = (SyscallDef.NumArgs + 1); Arg < FEXCore::HLE::SyscallArguments::MAX_ARGS; ++Arg) {
|
||||
IREmit->ReplaceNodeArgument(CodeNode, Arg, IREmit->Invalid());
|
||||
}
|
||||
#ifdef _M_ARM_64
|
||||
// Replace syscall with inline passthrough syscall if we can
|
||||
if (SyscallDef.HostSyscallNumber != -1) {
|
||||
IREmit->SetWriteCursor(CodeNode);
|
||||
// Skip Args[0] since that is the syscallid
|
||||
auto InlineSyscall =
|
||||
IREmit->_InlineSyscall(CurrentIR.GetNode(IROp->Args[1]), CurrentIR.GetNode(IROp->Args[2]), CurrentIR.GetNode(IROp->Args[3]),
|
||||
CurrentIR.GetNode(IROp->Args[4]), CurrentIR.GetNode(IROp->Args[5]), CurrentIR.GetNode(IROp->Args[6]),
|
||||
SyscallDef.HostSyscallNumber, Op->Flags);
|
||||
|
||||
// Replace all syscall uses with this inline one
|
||||
IREmit->ReplaceAllUsesWith(CodeNode, InlineSyscall);
|
||||
|
||||
// We must remove here since DCE can't remove a IROp with sideeffects
|
||||
IREmit->Remove(CodeNode);
|
||||
}
|
||||
#endif
|
||||
}
|
||||
|
||||
Changed = true;
|
||||
}
|
||||
} else if (IROp->Op == FEXCore::IR::OP_CPUID) {
|
||||
auto Op = IROp->CW<IR::IROp_CPUID>();
|
||||
|
||||
uint64_t ConstantFunction {}, ConstantLeaf {};
|
||||
bool IsConstantFunction = IREmit->IsValueConstant(Op->Function, &ConstantFunction);
|
||||
bool IsConstantLeaf = IREmit->IsValueConstant(Op->Leaf, &ConstantLeaf);
|
||||
// If the CPUID function is constant then we can try and optimize.
|
||||
if (IsConstantFunction) { // && ConstantFunction != 1) {
|
||||
// Check if it supports constant data reporting for this function.
|
||||
const auto SupportsConstant = CPUID->DoesFunctionReportConstantData(ConstantFunction);
|
||||
if (SupportsConstant.SupportsConstantFunction == CPUIDEmu::SupportsConstant::CONSTANT) {
|
||||
// If the CPUID needs a constant leaf to be optimized then this can't work if we didn't const-prop the leaf register.
|
||||
if (!(SupportsConstant.NeedsLeaf == CPUIDEmu::NeedsLeafConstant::NEEDSLEAFCONSTANT && !IsConstantLeaf)) {
|
||||
// Calculate the constant data and replace all uses.
|
||||
// DCE will remove the CPUID IR operation.
|
||||
const auto ConstantCPUIDResult = CPUID->RunFunction(ConstantFunction, ConstantLeaf);
|
||||
uint64_t ResultsLower = (static_cast<uint64_t>(ConstantCPUIDResult.ebx) << 32) | ConstantCPUIDResult.eax;
|
||||
uint64_t ResultsUpper = (static_cast<uint64_t>(ConstantCPUIDResult.edx) << 32) | ConstantCPUIDResult.ecx;
|
||||
IREmit->SetWriteCursor(CodeNode);
|
||||
auto ElementPair = IREmit->_CreateElementPair(IR::OpSize::i128Bit, IREmit->_Constant(ResultsLower), IREmit->_Constant(ResultsUpper));
|
||||
// Replace all CPUID uses with this inline one
|
||||
IREmit->ReplaceAllUsesWith(CodeNode, ElementPair);
|
||||
Changed = true;
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
else if (IROp->Op == FEXCore::IR::OP_XGETBV) {
|
||||
auto Op = IROp->CW<IR::IROp_XGetBV>();
|
||||
|
||||
uint64_t ConstantFunction {};
|
||||
if (IREmit->IsValueConstant(Op->Function, &ConstantFunction) && CPUID->DoesXCRFunctionReportConstantData(ConstantFunction)) {
|
||||
const auto ConstantXCRResult = CPUID->RunXCRFunction(ConstantFunction);
|
||||
IREmit->SetWriteCursor(CodeNode);
|
||||
auto ElementPair =
|
||||
IREmit->_CreateElementPair(IR::OpSize::i64Bit, IREmit->_Constant(ConstantXCRResult.eax), IREmit->_Constant(ConstantXCRResult.edx));
|
||||
// Replace all xgetbv uses with this inline one
|
||||
IREmit->ReplaceAllUsesWith(CodeNode, ElementPair);
|
||||
Changed = true;
|
||||
}
|
||||
}
|
||||
}
|
||||
return Changed;
|
||||
}
|
||||
|
||||
fextl::unique_ptr<FEXCore::IR::Pass> CreateInlineCallOptimization(const FEXCore::CPUIDEmu* CPUID) {
|
||||
return fextl::make_unique<InlineCallOptimization>(CPUID);
|
||||
}
|
||||
|
||||
} // namespace FEXCore::IR
|
||||
@@ -1,109 +0,0 @@
|
||||
// SPDX-License-Identifier: MIT
|
||||
/*
|
||||
$info$
|
||||
tags: ir|opts
|
||||
desc: Long divide elimination pass
|
||||
$end_info$
|
||||
*/
|
||||
|
||||
#include "Interface/IR/IREmitter.h"
|
||||
#include "Interface/IR/PassManager.h"
|
||||
#include <FEXCore/IR/IR.h>
|
||||
#include <FEXCore/Utils/Profiler.h>
|
||||
|
||||
#include <memory>
|
||||
#include <stdint.h>
|
||||
|
||||
namespace FEXCore::IR {
|
||||
|
||||
class LongDivideEliminationPass final : public FEXCore::IR::Pass {
|
||||
public:
|
||||
bool Run(IREmitter* IREmit) override;
|
||||
private:
|
||||
bool IsZeroOp(IREmitter* IREmit, OrderedNodeWrapper Arg);
|
||||
bool IsSextOp(IREmitter* IREmit, OrderedNodeWrapper Lower, OrderedNodeWrapper Upper);
|
||||
};
|
||||
|
||||
bool LongDivideEliminationPass::IsZeroOp(IREmitter* IREmit, OrderedNodeWrapper Arg) {
|
||||
uint64_t Value;
|
||||
|
||||
if (IREmit->IsValueConstant(Arg, &Value)) {
|
||||
// Zero constant based zero op
|
||||
return Value == 0;
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
bool LongDivideEliminationPass::IsSextOp(IREmitter* IREmit, OrderedNodeWrapper Lower, OrderedNodeWrapper Upper) {
|
||||
// We need to check if the upper source is a sext of the lower source
|
||||
auto UpperIROp = IREmit->GetOpHeader(Upper);
|
||||
if (UpperIROp->Op == OP_SBFE) {
|
||||
auto Op = UpperIROp->C<IR::IROp_Sbfe>();
|
||||
if (Op->Width == 1 && Op->lsb == 63) {
|
||||
// CQO: OrderedNode *Upper = _Sbfe(1, Size * 8 - 1, Src);
|
||||
// If the lower is the upper in this case then it can be optimized
|
||||
return Op->Header.Args[0] == Lower;
|
||||
}
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
bool LongDivideEliminationPass::Run(IREmitter* IREmit) {
|
||||
FEXCORE_PROFILE_SCOPED("PassManager::LDE");
|
||||
|
||||
bool Changed = false;
|
||||
auto CurrentIR = IREmit->ViewIR();
|
||||
auto OriginalWriteCursor = IREmit->GetWriteCursor();
|
||||
|
||||
for (auto [BlockNode, BlockHeader] : CurrentIR.GetBlocks()) {
|
||||
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
|
||||
if (IROp->Size == 8) {
|
||||
if (IROp->Op == OP_LDIV || IROp->Op == OP_LREM) {
|
||||
auto Op = IROp->C<IR::IROp_LDiv>();
|
||||
// Check upper Op to see if it came from a CQO
|
||||
// CQO: OrderedNode *Upper = _Sbfe(1, Size * 8 - 1, Src);
|
||||
// If it does then it we only need a 64bit SDIV
|
||||
if (IsSextOp(IREmit, Op->Lower, Op->Upper)) {
|
||||
IREmit->SetWriteCursor(CodeNode);
|
||||
OrderedNode* Lower = CurrentIR.GetNode(Op->Lower);
|
||||
OrderedNode* Divisor = CurrentIR.GetNode(Op->Divisor);
|
||||
OrderedNode* SDivOp {};
|
||||
if (IROp->Op == OP_LDIV) {
|
||||
SDivOp = IREmit->_Div(OpSize::i64Bit, Lower, Divisor);
|
||||
} else {
|
||||
SDivOp = IREmit->_Rem(OpSize::i64Bit, Lower, Divisor);
|
||||
}
|
||||
IREmit->ReplaceAllUsesWith(CodeNode, SDivOp);
|
||||
Changed = true;
|
||||
}
|
||||
} else if (IROp->Op == OP_LUDIV || IROp->Op == OP_LUREM) {
|
||||
auto Op = IROp->C<IR::IROp_LUDiv>();
|
||||
// Check upper Op to see if it came from a zeroing op
|
||||
// If it does then it we only need a 64bit UDIV
|
||||
if (IsZeroOp(IREmit, Op->Upper)) {
|
||||
IREmit->SetWriteCursor(CodeNode);
|
||||
OrderedNode* Lower = CurrentIR.GetNode(Op->Lower);
|
||||
OrderedNode* Divisor = CurrentIR.GetNode(Op->Divisor);
|
||||
OrderedNode* UDivOp {};
|
||||
if (IROp->Op == OP_LUDIV) {
|
||||
UDivOp = IREmit->_UDiv(OpSize::i64Bit, Lower, Divisor);
|
||||
} else {
|
||||
UDivOp = IREmit->_URem(OpSize::i64Bit, Lower, Divisor);
|
||||
}
|
||||
IREmit->ReplaceAllUsesWith(CodeNode, UDivOp);
|
||||
Changed = true;
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
IREmit->SetWriteCursor(OriginalWriteCursor);
|
||||
|
||||
return Changed;
|
||||
}
|
||||
|
||||
fextl::unique_ptr<FEXCore::IR::Pass> CreateLongDivideEliminationPass() {
|
||||
return fextl::make_unique<LongDivideEliminationPass>();
|
||||
}
|
||||
} // namespace FEXCore::IR
|
||||
@@ -23,14 +23,12 @@ struct RegState {
|
||||
static constexpr IR::NodeID UninitializedValue {0};
|
||||
static constexpr IR::NodeID InvalidReg {0xffff'ffff};
|
||||
static constexpr IR::NodeID CorruptedPair {0xffff'fffe};
|
||||
static constexpr IR::NodeID ClobberedValue {0xffff'fffd};
|
||||
static constexpr IR::NodeID StaticAssigned {0xffff'ff00};
|
||||
|
||||
// This class makes some assumptions about how the host registers are arranged and mapped to virtual registers:
|
||||
// 1. There will be less than 32 GPRs and 32 FPRs
|
||||
// 2. If the GPRFixed class is used, there will be 16 GPRs and 16 FixedGPRs max
|
||||
// 3. Same with FPRFixed
|
||||
// 4. If the GPRPairClass is used, it is assumed each GPRPair N will map onto GPRs N*2 and N*2 + 1
|
||||
// 4. If the GPRPairClass is used, it is assumed each GPRPair N will map onto GPRs N and N + 1
|
||||
|
||||
// These assumptions were all true for the state of the arm64 and x86 jits at the time this was written
|
||||
|
||||
@@ -53,8 +51,8 @@ struct RegState {
|
||||
return true;
|
||||
case GPRPairClass:
|
||||
// Alias paired registers onto both
|
||||
GPRs[Reg.Reg * 2] = ssa;
|
||||
GPRs[Reg.Reg * 2 + 1] = ssa;
|
||||
GPRs[Reg.Reg] = ssa;
|
||||
GPRs[Reg.Reg + 1] = ssa;
|
||||
return true;
|
||||
}
|
||||
return false;
|
||||
@@ -70,8 +68,8 @@ struct RegState {
|
||||
case FPRFixedClass: return FPRsFixed[Reg.Reg];
|
||||
case GPRPairClass:
|
||||
// Make sure both halves of the Pair contain the same SSA
|
||||
if (GPRs[Reg.Reg * 2] == GPRs[Reg.Reg * 2 + 1]) {
|
||||
return GPRs[Reg.Reg * 2];
|
||||
if (GPRs[Reg.Reg] == GPRs[Reg.Reg + 1]) {
|
||||
return GPRs[Reg.Reg];
|
||||
}
|
||||
return CorruptedPair;
|
||||
}
|
||||
@@ -83,91 +81,12 @@ struct RegState {
|
||||
Spills[SpillSlot] = ssa;
|
||||
}
|
||||
|
||||
// Consume (and return) the SSA id currently in a spill slot
|
||||
// Return the SSA id currently in a spill slot
|
||||
IR::NodeID Unspill(uint32_t SpillSlot) {
|
||||
if (Spills.contains(SpillSlot)) {
|
||||
const auto Value = Spills[SpillSlot];
|
||||
Spills.erase(SpillSlot);
|
||||
return Value;
|
||||
}
|
||||
return UninitializedValue;
|
||||
}
|
||||
|
||||
// Intersect another regstate with this one
|
||||
// Any registers/slots which contain the same SSA id will be persevered
|
||||
// Anything else will be marked as Clobbered
|
||||
//
|
||||
// Useful for merging two branches of control flow.
|
||||
// Any register that differs depending on control flow shouldn't be consumed by
|
||||
// code that follows
|
||||
void Intersect(RegState& other) {
|
||||
for (size_t i = 0; i < GPRs.size(); i++) {
|
||||
if (GPRs[i] != other.GPRs[i]) {
|
||||
GPRs[i] = ClobberedValue;
|
||||
}
|
||||
}
|
||||
|
||||
for (size_t i = 0; i < GPRsFixed.size(); i++) {
|
||||
if (GPRsFixed[i] != other.GPRsFixed[i]) {
|
||||
GPRsFixed[i] = ClobberedValue;
|
||||
}
|
||||
}
|
||||
|
||||
for (size_t i = 0; i < FPRs.size(); i++) {
|
||||
if (FPRs[i] != other.FPRs[i]) {
|
||||
FPRs[i] = ClobberedValue;
|
||||
}
|
||||
}
|
||||
|
||||
for (size_t i = 0; i < FPRsFixed.size(); i++) {
|
||||
if (FPRsFixed[i] != other.FPRsFixed[i]) {
|
||||
FPRsFixed[i] = ClobberedValue;
|
||||
}
|
||||
}
|
||||
|
||||
for (auto it = Spills.begin(); it != Spills.end(); it++) {
|
||||
auto& [SlotID, Value] = *it;
|
||||
if (!other.Spills.contains(SlotID)) {
|
||||
Spills.erase(it);
|
||||
} else if (Value != other.Spills[SlotID]) {
|
||||
Value = ClobberedValue;
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Filter out all registers/slots containing an SSA id larger than MaxSSA
|
||||
// Mark them as Clobbered.
|
||||
// Useful for backwards edges, where using an SSA from before the
|
||||
void Filter(IR::NodeID MaxSSA) {
|
||||
for (auto& gpr : GPRs) {
|
||||
if (gpr > MaxSSA) {
|
||||
gpr = ClobberedValue;
|
||||
}
|
||||
}
|
||||
|
||||
for (auto& gpr : GPRsFixed) {
|
||||
if (gpr > MaxSSA) {
|
||||
gpr = ClobberedValue;
|
||||
}
|
||||
}
|
||||
|
||||
for (auto& fpr : FPRs) {
|
||||
if (fpr > MaxSSA) {
|
||||
fpr = ClobberedValue;
|
||||
}
|
||||
}
|
||||
|
||||
for (auto& fpr : FPRsFixed) {
|
||||
if (fpr > MaxSSA) {
|
||||
fpr = ClobberedValue;
|
||||
}
|
||||
}
|
||||
|
||||
for (auto it = Spills.begin(); it != Spills.end(); it++) {
|
||||
auto& [SlotID, Value] = *it;
|
||||
if (Value > MaxSSA) {
|
||||
Spills.erase(it);
|
||||
}
|
||||
return Spills[SpillSlot];
|
||||
} else {
|
||||
return UninitializedValue;
|
||||
}
|
||||
}
|
||||
|
||||
@@ -178,119 +97,32 @@ private:
|
||||
std::array<IR::NodeID, 32> FPRs = {};
|
||||
|
||||
fextl::unordered_map<uint32_t, IR::NodeID> Spills;
|
||||
|
||||
public:
|
||||
uint32_t Version {}; // Used to force regeneration of RegStates after following backward edges
|
||||
};
|
||||
|
||||
class RAValidation final : public FEXCore::IR::Pass {
|
||||
public:
|
||||
~RAValidation() {}
|
||||
bool Run(IREmitter* IREmit) override;
|
||||
|
||||
private:
|
||||
// Holds the calculated RegState at the exit of each block
|
||||
fextl::unordered_map<IR::NodeID, RegState> BlockExitState;
|
||||
|
||||
// A queue of blocks we need to visit (or revisit)
|
||||
fextl::deque<OrderedNode*> BlocksToVisit;
|
||||
void Run(IREmitter* IREmit) override;
|
||||
};
|
||||
|
||||
|
||||
bool RAValidation::Run(IREmitter* IREmit) {
|
||||
void RAValidation::Run(IREmitter* IREmit) {
|
||||
if (!Manager->HasPass("RA")) {
|
||||
return false;
|
||||
return;
|
||||
}
|
||||
|
||||
FEXCORE_PROFILE_SCOPED("PassManager::RAValidation");
|
||||
|
||||
IR::RegisterAllocationData* RAData = Manager->GetPass<IR::RegisterAllocationPass>("RA")->GetAllocationData();
|
||||
BlockExitState.clear();
|
||||
// BlocksToVisit will already be empty
|
||||
|
||||
// Get the control flow graph from the validation pass
|
||||
auto ValidationPass = Manager->GetPass<IRValidation>("IRValidation");
|
||||
LOGMAN_THROW_AA_FMT(ValidationPass != nullptr, "Couldn't find IRValidation pass");
|
||||
|
||||
auto& OffsetToBlockMap = ValidationPass->OffsetToBlockMap;
|
||||
|
||||
LOGMAN_THROW_AA_FMT(ValidationPass->EntryBlock != nullptr, "No entry point");
|
||||
BlocksToVisit.push_front(ValidationPass->EntryBlock); // Currently only a single entry point
|
||||
|
||||
bool HadError = false;
|
||||
fextl::ostringstream Errors;
|
||||
|
||||
auto CurrentIR = IREmit->ViewIR();
|
||||
uint32_t CurrentVersion = 1; // Incremented every backwards edge
|
||||
|
||||
while (!BlocksToVisit.empty()) {
|
||||
auto BlockNode = BlocksToVisit.front();
|
||||
const auto BlockID = CurrentIR.GetID(BlockNode);
|
||||
auto& BlockInfo = OffsetToBlockMap[BlockID];
|
||||
|
||||
const auto IsFowardsEdge = [&](IR::NodeID PredecessorID) {
|
||||
// Blocks are sorted in FEXes IR, so backwards edges always go to a lower (or equal) Block ID
|
||||
return PredecessorID < BlockID;
|
||||
};
|
||||
|
||||
// First, make sure we have the exit state for all Predecessors that
|
||||
// get here via a forwards branch.
|
||||
bool MissingPredecessor = false;
|
||||
|
||||
for (auto Predecessor : BlockInfo.Predecessors) {
|
||||
const auto PredecessorID = CurrentIR.GetID(Predecessor);
|
||||
const bool HaveState = BlockExitState.contains(PredecessorID) && BlockExitState[PredecessorID].Version == CurrentVersion;
|
||||
|
||||
if (IsFowardsEdge(PredecessorID) && !HaveState) {
|
||||
// We are probably about to visit this node anyway, remove it
|
||||
std::remove(BlocksToVisit.begin(), BlocksToVisit.end(), Predecessor);
|
||||
|
||||
// Add the missing predecessor to start of queue
|
||||
BlocksToVisit.push_front(Predecessor);
|
||||
MissingPredecessor = true;
|
||||
}
|
||||
}
|
||||
|
||||
if (MissingPredecessor) {
|
||||
// We'll have to come back to this block later
|
||||
continue;
|
||||
}
|
||||
|
||||
// We have committed to processing this block
|
||||
// Remove from queue
|
||||
BlocksToVisit.pop_front();
|
||||
|
||||
bool FirstVisit = !BlockExitState.contains(BlockID);
|
||||
|
||||
// Second, we need to determine the register status as of Block entry
|
||||
auto BlockOp = CurrentIR.GetOp<IROp_CodeBlock>(BlockNode);
|
||||
const auto FirstSSA = BlockOp->Begin.ID();
|
||||
|
||||
auto& BlockRegState = BlockExitState.try_emplace(BlockID).first->second;
|
||||
bool EmptyRegState = true;
|
||||
auto Intersect = [&](RegState& Other) {
|
||||
if (EmptyRegState) {
|
||||
BlockRegState = Other;
|
||||
EmptyRegState = false;
|
||||
} else {
|
||||
BlockRegState.Intersect(Other);
|
||||
}
|
||||
};
|
||||
|
||||
for (auto Predecessor : BlockInfo.Predecessors) {
|
||||
auto PredecessorID = CurrentIR.GetID(Predecessor);
|
||||
if (BlockExitState.contains(PredecessorID)) {
|
||||
if (IsFowardsEdge(PredecessorID)) {
|
||||
Intersect(BlockExitState[PredecessorID]);
|
||||
} else {
|
||||
RegState Filtered = BlockExitState[PredecessorID];
|
||||
Filtered.Filter(FirstSSA);
|
||||
Intersect(Filtered);
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Third, we need to iterate over all IR ops in the block
|
||||
for (auto [BlockNode, BlockIROp] : CurrentIR.GetBlocks()) {
|
||||
// We only allocate registers locally, so state is reset each block
|
||||
struct RegState BlockRegState = {};
|
||||
|
||||
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
|
||||
const auto ID = CurrentIR.GetID(CodeNode);
|
||||
@@ -319,11 +151,6 @@ bool RAValidation::Run(IREmitter* IREmit) {
|
||||
HadError |= true;
|
||||
|
||||
Errors << fextl::fmt::format("%{}: Arg[{}] expects reg{} to contain %{}, but it is uninitialized\n", ID, i, PhyReg.Reg, ArgID);
|
||||
} else if (CurrentSSAAtReg == RegState::ClobberedValue) {
|
||||
HadError |= true;
|
||||
|
||||
Errors << fextl::fmt::format("%{}: Arg[{}] expects reg{} to contain %{}, but contents vary depending on control flow\n", ID, i,
|
||||
PhyReg.Reg, ArgID);
|
||||
} else if (CurrentSSAAtReg != ArgID) {
|
||||
HadError |= true;
|
||||
Errors << fextl::fmt::format("%{}: Arg[{}] expects reg{} to contain %{}, but it actually contains %{}\n", ID, i, PhyReg.Reg,
|
||||
@@ -346,20 +173,13 @@ bool RAValidation::Run(IREmitter* IREmit) {
|
||||
const auto Value = BlockRegState.Unspill(FillRegister->Slot);
|
||||
|
||||
// TODO: This only proves that the Spill has a consistent SSA value
|
||||
// In the future we need to prove it contains the correct SSA value
|
||||
// In the future we need to prove it contains the correct SSA value. For
|
||||
// this we need to analyze copies/swaps properly. As a hot fix, don't
|
||||
// compare Value with ExpectedValue.
|
||||
|
||||
if (Value == RegState::UninitializedValue) {
|
||||
HadError |= true;
|
||||
Errors << fextl::fmt::format("%{}: FillRegister expected %{} in Slot {}, but was undefined in at least one control flow path\n",
|
||||
ID, ExpectedValue, FillRegister->Slot);
|
||||
} else if (Value == RegState::ClobberedValue) {
|
||||
HadError |= true;
|
||||
Errors << fextl::fmt::format("%{}: FillRegister expected %{} in Slot {}, but contents vary depending on control flow\n", ID,
|
||||
ExpectedValue, FillRegister->Slot);
|
||||
} else if (Value != ExpectedValue) {
|
||||
HadError |= true;
|
||||
Errors << fextl::fmt::format("%{}: FillRegister expected %{} in Slot {}, but it actually contains %{}\n", ID, ExpectedValue,
|
||||
FillRegister->Slot, Value);
|
||||
Errors << fextl::fmt::format("%{}: FillRegister expected %{} in Slot {}, but was undefined\n", ID, ExpectedValue, FillRegister->Slot);
|
||||
}
|
||||
break;
|
||||
}
|
||||
@@ -375,74 +195,10 @@ bool RAValidation::Run(IREmitter* IREmit) {
|
||||
}
|
||||
|
||||
// Update BlockState map
|
||||
BlockRegState.Set(RAData->GetNodeRegister(ID), ID);
|
||||
}
|
||||
|
||||
// Forth, Add successors to the queue of blocks to validate
|
||||
for (auto Successor : BlockInfo.Successors) {
|
||||
auto SuccessorID = CurrentIR.GetID(Successor);
|
||||
|
||||
// Blocks are sorted in FEXes IR, so backwards edges always go to a lower (or equal) Block ID
|
||||
bool FowardsEdge = SuccessorID > BlockID;
|
||||
|
||||
if (FowardsEdge) {
|
||||
// Always follow forwards edges, assuming it's not already on the queue
|
||||
if (std::find(BlocksToVisit.begin(), BlocksToVisit.end(), Successor) == std::end(BlocksToVisit)) {
|
||||
// Push to the back of queue so there is a higher chance all predecessors for this block are done first
|
||||
BlocksToVisit.push_back(Successor);
|
||||
}
|
||||
} else if (FirstVisit) {
|
||||
// Now that we have the block data for the backwards edge, we can visit it again and make
|
||||
// sure it (and all it's successors) are still valid.
|
||||
|
||||
// But only do this the first time we encounter each backwards edge.
|
||||
|
||||
// Push to the front of queue, so we get this re-checking done before examining future nodes.
|
||||
BlocksToVisit.push_front(Successor);
|
||||
|
||||
// Make sure states are reprocessed
|
||||
CurrentVersion++;
|
||||
if (IROp->Op != OP_SPILLREGISTER) {
|
||||
BlockRegState.Set(RAData->GetNodeRegister(ID), ID);
|
||||
}
|
||||
}
|
||||
|
||||
BlockRegState.Version = CurrentVersion;
|
||||
|
||||
if (CurrentVersion > 10000) {
|
||||
Errors << "Infinite Loop\n";
|
||||
HadError |= true;
|
||||
|
||||
for (auto [BlockNode, BlockHeader] : CurrentIR.GetBlocks()) {
|
||||
const auto BlockID = CurrentIR.GetID(BlockNode);
|
||||
const auto& BlockInfo = OffsetToBlockMap[BlockID];
|
||||
|
||||
Errors << fextl::fmt::format("Block {}\n\tPredecessors: ", BlockID);
|
||||
|
||||
for (auto Predecessor : BlockInfo.Predecessors) {
|
||||
const auto PredecessorID = CurrentIR.GetID(Predecessor);
|
||||
const bool FowardsEdge = PredecessorID < BlockID;
|
||||
if (!FowardsEdge) {
|
||||
Errors << "(Backwards): ";
|
||||
}
|
||||
Errors << fextl::fmt::format("Block {} ", PredecessorID);
|
||||
}
|
||||
|
||||
Errors << "\n\tSuccessors: ";
|
||||
|
||||
for (auto Successor : BlockInfo.Successors) {
|
||||
const auto SuccessorID = CurrentIR.GetID(Successor);
|
||||
const bool FowardsEdge = SuccessorID > BlockID;
|
||||
|
||||
if (!FowardsEdge) {
|
||||
Errors << "(Backwards): ";
|
||||
}
|
||||
Errors << fextl::fmt::format("Block {} ", SuccessorID);
|
||||
}
|
||||
|
||||
Errors << "\n\n";
|
||||
}
|
||||
|
||||
break;
|
||||
}
|
||||
}
|
||||
|
||||
if (HadError) {
|
||||
@@ -454,8 +210,6 @@ bool RAValidation::Run(IREmitter* IREmit) {
|
||||
|
||||
Errors.clear();
|
||||
}
|
||||
|
||||
return false;
|
||||
}
|
||||
|
||||
fextl::unique_ptr<FEXCore::IR::Pass> CreateRAValidation() {
|
||||
|
||||
@@ -53,16 +53,17 @@ struct FlagInfo {
|
||||
|
||||
class DeadFlagCalculationEliminination final : public FEXCore::IR::Pass {
|
||||
public:
|
||||
bool Run(IREmitter* IREmit) override;
|
||||
void Run(IREmitter* IREmit) override;
|
||||
|
||||
private:
|
||||
FlagInfo Classify(IROp_Header* Node);
|
||||
unsigned FlagForOffset(unsigned Offset);
|
||||
unsigned FlagForReg(unsigned Reg);
|
||||
unsigned FlagsForCondClassType(CondClassType Cond);
|
||||
bool EliminateDeadCode(IREmitter* IREmit, Ref CodeNode, IROp_Header* IROp);
|
||||
};
|
||||
|
||||
unsigned DeadFlagCalculationEliminination::FlagForOffset(unsigned Offset) {
|
||||
return Offset == offsetof(FEXCore::Core::CPUState, pf_raw) ? FLAG_P : Offset == offsetof(FEXCore::Core::CPUState, af_raw) ? FLAG_A : 0;
|
||||
unsigned DeadFlagCalculationEliminination::FlagForReg(unsigned Reg) {
|
||||
return Reg == Core::CPUState::PF_AS_GREG ? FLAG_P : Reg == Core::CPUState::AF_AS_GREG ? FLAG_A : 0;
|
||||
};
|
||||
|
||||
unsigned DeadFlagCalculationEliminination::FlagsForCondClassType(CondClassType Cond) {
|
||||
@@ -283,21 +284,20 @@ FlagInfo DeadFlagCalculationEliminination::Classify(IROp_Header* IROp) {
|
||||
|
||||
case OP_LOADREGISTER: {
|
||||
auto Op = IROp->CW<IR::IROp_LoadRegister>();
|
||||
if (Op->Class != GPRClass || Op->StaticClass != GPRFixedClass) {
|
||||
if (Op->Class != GPRClass) {
|
||||
break;
|
||||
}
|
||||
|
||||
return {.Read = FlagForOffset(Op->Offset)};
|
||||
return {.Read = FlagForReg(Op->Reg)};
|
||||
}
|
||||
|
||||
case OP_STOREREGISTER: {
|
||||
auto Op = IROp->CW<IR::IROp_StoreRegister>();
|
||||
if (Op->Class != GPRClass || Op->StaticClass != GPRFixedClass) {
|
||||
if (Op->Class != GPRClass) {
|
||||
break;
|
||||
}
|
||||
|
||||
LOGMAN_THROW_A_FMT(!Op->IsPrewrite, "PF/AF writes are fixed-form");
|
||||
unsigned Flag = FlagForOffset(Op->Offset);
|
||||
unsigned Flag = FlagForReg(Op->Reg);
|
||||
|
||||
return {
|
||||
.Write = Flag,
|
||||
@@ -311,13 +311,53 @@ FlagInfo DeadFlagCalculationEliminination::Classify(IROp_Header* IROp) {
|
||||
return {.Trivial = true};
|
||||
}
|
||||
|
||||
// General purpose dead code elimination. Returns whether flag handling should
|
||||
// be skipped (because it was removed or could not possibly affect flags).
|
||||
bool DeadFlagCalculationEliminination::EliminateDeadCode(IREmitter* IREmit, Ref CodeNode, IROp_Header* IROp) {
|
||||
// Can't remove anything used or with side effects.
|
||||
if (CodeNode->GetUses() > 0 || IR::HasSideEffects(IROp->Op)) {
|
||||
return false;
|
||||
}
|
||||
|
||||
switch (IROp->Op) {
|
||||
case OP_SYSCALL: {
|
||||
auto Op = IROp->C<IR::IROp_Syscall>();
|
||||
if ((Op->Flags & IR::SyscallFlags::NOSIDEEFFECTS) != IR::SyscallFlags::NOSIDEEFFECTS) {
|
||||
return false;
|
||||
}
|
||||
|
||||
break;
|
||||
}
|
||||
case OP_INLINESYSCALL: {
|
||||
auto Op = IROp->C<IR::IROp_Syscall>();
|
||||
if ((Op->Flags & IR::SyscallFlags::NOSIDEEFFECTS) != IR::SyscallFlags::NOSIDEEFFECTS) {
|
||||
return false;
|
||||
}
|
||||
|
||||
break;
|
||||
}
|
||||
|
||||
// If the result of the atomic fetch is completely unused, convert it to a non-fetching atomic operation.
|
||||
case OP_ATOMICFETCHADD: IROp->Op = OP_ATOMICADD; return true;
|
||||
case OP_ATOMICFETCHSUB: IROp->Op = OP_ATOMICSUB; return true;
|
||||
case OP_ATOMICFETCHAND: IROp->Op = OP_ATOMICAND; return true;
|
||||
case OP_ATOMICFETCHCLR: IROp->Op = OP_ATOMICCLR; return true;
|
||||
case OP_ATOMICFETCHOR: IROp->Op = OP_ATOMICOR; return true;
|
||||
case OP_ATOMICFETCHXOR: IROp->Op = OP_ATOMICXOR; return true;
|
||||
case OP_ATOMICFETCHNEG: IROp->Op = OP_ATOMICNEG; return true;
|
||||
default: break;
|
||||
}
|
||||
|
||||
IREmit->Remove(CodeNode);
|
||||
return true;
|
||||
}
|
||||
|
||||
/**
|
||||
* @brief This pass removes flag calculations that will otherwise be unused INSIDE of that block
|
||||
* @brief This pass removes dead code locally.
|
||||
*/
|
||||
bool DeadFlagCalculationEliminination::Run(IREmitter* IREmit) {
|
||||
void DeadFlagCalculationEliminination::Run(IREmitter* IREmit) {
|
||||
FEXCORE_PROFILE_SCOPED("PassManager::DFE");
|
||||
|
||||
bool Changed = false;
|
||||
auto CurrentIR = IREmit->ViewIR();
|
||||
|
||||
for (auto [BlockNode, BlockHeader] : CurrentIR.GetBlocks()) {
|
||||
@@ -339,13 +379,7 @@ bool DeadFlagCalculationEliminination::Run(IREmitter* IREmit) {
|
||||
// Optimizing flags can cause earlier flag reads to become dead but dead
|
||||
// flag reads should not impede optimiation of earlier dead flag writes.
|
||||
// We must DCE as we go to ensure we converge in a single iteration.
|
||||
//
|
||||
// TODO: This whole pass could be merged with DCE?
|
||||
bool HasSideEffects = IR::HasSideEffects(IROp->Op);
|
||||
if (!HasSideEffects && CodeNode->GetUses() == 0) {
|
||||
Changed = true;
|
||||
IREmit->Remove(CodeNode);
|
||||
} else {
|
||||
if (!EliminateDeadCode(IREmit, CodeNode, IROp)) {
|
||||
// Optimiation algorithm: For each flag written...
|
||||
//
|
||||
// If the flag has a later read (per FlagsRead), remove the flag from
|
||||
@@ -365,13 +399,11 @@ bool DeadFlagCalculationEliminination::Run(IREmitter* IREmit) {
|
||||
bool Eliminated = false;
|
||||
|
||||
if ((FlagsRead & Info.Write) == 0) {
|
||||
if (Info.CanEliminate && CodeNode->GetUses() == 0) {
|
||||
if ((Info.CanEliminate || Info.CanReplace) && CodeNode->GetUses() == 0) {
|
||||
IREmit->Remove(CodeNode);
|
||||
Eliminated = true;
|
||||
Changed = true;
|
||||
} else if (Info.CanReplace) {
|
||||
IROp->Op = Info.Replacement;
|
||||
Changed = true;
|
||||
}
|
||||
} else {
|
||||
FlagsRead &= ~Info.Write;
|
||||
@@ -393,8 +425,6 @@ bool DeadFlagCalculationEliminination::Run(IREmitter* IREmit) {
|
||||
--CodeLast;
|
||||
}
|
||||
}
|
||||
|
||||
return Changed;
|
||||
}
|
||||
|
||||
fextl::unique_ptr<FEXCore::IR::Pass> CreateDeadFlagCalculationEliminination() {
|
||||
|
||||
File diff suppressed because it is too large.
Load diff
@@ -18,31 +18,10 @@ struct RegisterClassType;
|
||||
|
||||
class RegisterAllocationPass : public FEXCore::IR::Pass {
|
||||
public:
|
||||
bool HasFullRA() const {
|
||||
return HadFullRA;
|
||||
}
|
||||
|
||||
virtual void AllocateRegisterSet(uint32_t ClassCount) = 0;
|
||||
virtual void AddRegisters(FEXCore::IR::RegisterClassType Class, uint32_t RegisterCount) = 0;
|
||||
|
||||
/**
|
||||
* @brief Adds a conflict between the two registers (and their respective classes) so if one is allocated then the other can not be
|
||||
* allocated in the same live range.
|
||||
*
|
||||
* Conflict is added both directions, so only necessary to add a conflict one way
|
||||
*
|
||||
* ex:
|
||||
* AddRegisters(GPRClass, 2); -> {x0, x1} added to register class GPR
|
||||
* AddRegisters(PairClass, 1); -> {{x0, x1}} pair added to class GPR
|
||||
* AddRegisterConflict(GPRClass, 0, PairClass, 0); -> Make sure the pair interferes with x0
|
||||
* AddRegisterConflict(GPRClass, 1, PairClass, 1); -> Make sure the pair interferes with x1
|
||||
*/
|
||||
virtual void AddRegisterConflict(FEXCore::IR::RegisterClassType ClassConflict, uint32_t RegConflict, FEXCore::IR::RegisterClassType Class,
|
||||
uint32_t Reg) = 0;
|
||||
|
||||
/**
|
||||
* @name Inference graph handling
|
||||
* @{ */
|
||||
// Number of GPRs usable for pairs at start of GPR set. Must be even.
|
||||
uint32_t PairRegs;
|
||||
|
||||
/**
|
||||
* @brief Returns the register and class map array
|
||||
@@ -53,15 +32,6 @@ public:
|
||||
* @brief Returns and transfers ownership of the register and class map array
|
||||
*/
|
||||
virtual std::unique_ptr<RegisterAllocationData, RegisterAllocationDataDeleter> PullAllocationData() = 0;
|
||||
/** @} */
|
||||
|
||||
protected:
|
||||
bool HasSpills {};
|
||||
// Debug option to disable split slot reuse
|
||||
// Can be useful for testing if there is a bug with spill slots
|
||||
constexpr static bool ReuseSpillSlots {true};
|
||||
uint32_t SpillSlotCount {};
|
||||
bool HadFullRA {};
|
||||
};
|
||||
|
||||
} // namespace FEXCore::IR
|
||||
@@ -1,106 +0,0 @@
|
||||
// SPDX-License-Identifier: MIT
|
||||
/*
|
||||
$info$
|
||||
tags: ir|opts
|
||||
desc: Sanity Checking
|
||||
$end_info$
|
||||
*/
|
||||
|
||||
#include "Interface/IR/IR.h"
|
||||
#include "Interface/IR/IREmitter.h"
|
||||
#include "Interface/IR/PassManager.h"
|
||||
|
||||
#include <FEXCore/IR/IR.h>
|
||||
#include <FEXCore/Utils/LogManager.h>
|
||||
#include <FEXCore/Utils/Profiler.h>
|
||||
#include <FEXCore/fextl/set.h>
|
||||
#include <FEXCore/fextl/sstream.h>
|
||||
#include <FEXCore/fextl/unordered_map.h>
|
||||
#include <FEXCore/fextl/vector.h>
|
||||
|
||||
#include <functional>
|
||||
#include <memory>
|
||||
#include <stdint.h>
|
||||
#include <utility>
|
||||
|
||||
namespace FEXCore::IR::Validation {
|
||||
class ValueDominanceValidation final : public FEXCore::IR::Pass {
|
||||
public:
|
||||
bool Run(IREmitter* IREmit) override;
|
||||
};
|
||||
|
||||
bool ValueDominanceValidation::Run(IREmitter* IREmit) {
|
||||
FEXCORE_PROFILE_SCOPED("PassManager::ValueDominanceValidation");
|
||||
|
||||
bool HadError = false;
|
||||
auto CurrentIR = IREmit->ViewIR();
|
||||
|
||||
fextl::ostringstream Errors;
|
||||
|
||||
for (auto [BlockNode, BlockHeader] : CurrentIR.GetBlocks()) {
|
||||
auto BlockIROp = BlockHeader->CW<FEXCore::IR::IROp_CodeBlock>();
|
||||
|
||||
for (auto [CodeNode, IROp] : CurrentIR.GetCode(BlockNode)) {
|
||||
const auto CodeID = CurrentIR.GetID(CodeNode);
|
||||
|
||||
const uint8_t NumArgs = IR::GetRAArgs(IROp->Op);
|
||||
for (uint32_t i = 0; i < NumArgs; ++i) {
|
||||
if (IROp->Args[i].IsInvalid()) {
|
||||
continue;
|
||||
}
|
||||
|
||||
// We do not validate the location of inline constants because it's
|
||||
// irrelevant, they're ignored by RA and always inlined to where they
|
||||
// need to be. This lets us pool inline constants globally.
|
||||
IROps Op = CurrentIR.GetOp<IROp_Header>(IROp->Args[i])->Op;
|
||||
if (Op == OP_IRHEADER || Op == OP_INLINECONSTANT) {
|
||||
continue;
|
||||
}
|
||||
|
||||
OrderedNodeWrapper Arg = IROp->Args[i];
|
||||
|
||||
// If the SSA argument is not defined INSIDE the block, we have
|
||||
// cross-block liveness, which we forbid in the IR to simplify RA.
|
||||
if (!(Arg.ID() >= BlockIROp->Begin.ID() && Arg.ID() < BlockIROp->Last.ID())) {
|
||||
HadError |= true;
|
||||
Errors << "Inst %" << CodeID << ": Arg[" << i << "] %" << Arg.ID() << " definition not local!" << std::endl;
|
||||
continue;
|
||||
}
|
||||
|
||||
// The SSA argument is defined INSIDE this block.
|
||||
// It must only be declared prior to this instruction
|
||||
// Eg: Valid
|
||||
// CodeBlock_1:
|
||||
// %_1 = Load
|
||||
// %_2 = Load
|
||||
// %_3 = <Op> %_1, %_2
|
||||
//
|
||||
// Eg: Invalid
|
||||
// CodeBlock_1:
|
||||
// %_1 = Load
|
||||
// %_2 = <Op> %_1, %_3
|
||||
// %_3 = Load
|
||||
if (Arg.ID() > CodeID) {
|
||||
HadError |= true;
|
||||
Errors << "Inst %" << CodeID << ": Arg[" << i << "] %" << Arg.ID() << " definition does not dominate this use!" << std::endl;
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
if (HadError) {
|
||||
fextl::stringstream Out;
|
||||
FEXCore::IR::Dump(&Out, &CurrentIR, nullptr);
|
||||
Out << "Errors:" << std::endl << Errors.str() << std::endl;
|
||||
LogMan::Msg::EFmt("{}", Out.str());
|
||||
LOGMAN_MSG_A_FMT("Encountered IR validation Error");
|
||||
}
|
||||
|
||||
return false;
|
||||
}
|
||||
|
||||
fextl::unique_ptr<FEXCore::IR::Pass> CreateValueDominanceValidation() {
|
||||
return fextl::make_unique<ValueDominanceValidation>();
|
||||
}
|
||||
|
||||
} // namespace FEXCore::IR::Validation
|
||||
@@ -43,7 +43,6 @@ class FEX_PACKED RegisterAllocationData {
|
||||
public:
|
||||
uint32_t SpillSlotCount {};
|
||||
uint32_t MapCount {};
|
||||
bool IsShared {false};
|
||||
PhysicalRegister Map[0];
|
||||
|
||||
PhysicalRegister GetNodeRegister(NodeID Node) const {
|
||||
@@ -67,18 +66,13 @@ public:
|
||||
stream.Write((const char*)&SpillSlotCount, sizeof(SpillSlotCount));
|
||||
stream.Write((const char*)&MapCount, sizeof(MapCount));
|
||||
// RAData (inline)
|
||||
// In file, IsShared is always set
|
||||
bool _IsShared = true;
|
||||
stream.Write((const char*)&_IsShared, sizeof(IsShared));
|
||||
stream.Write((const char*)&Map[0], sizeof(Map[0]) * MapCount);
|
||||
}
|
||||
};
|
||||
|
||||
struct RegisterAllocationDataDeleter {
|
||||
void operator()(RegisterAllocationData* r) const {
|
||||
if (!r->IsShared) {
|
||||
FEXCore::Allocator::free(r);
|
||||
}
|
||||
FEXCore::Allocator::free(r);
|
||||
}
|
||||
};
|
||||
|
||||
@@ -87,7 +81,6 @@ inline auto RegisterAllocationData::Create(uint32_t NodeCount) -> UniquePtr {
|
||||
memset(&Ret->Map[0], PhysicalRegister::Invalid().Raw, NodeCount);
|
||||
Ret->SpillSlotCount = 0;
|
||||
Ret->MapCount = NodeCount;
|
||||
Ret->IsShared = false;
|
||||
return UniquePtr {Ret};
|
||||
}
|
||||
|
||||
@@ -96,7 +89,6 @@ inline auto RegisterAllocationData::CreateCopy() const -> UniquePtr {
|
||||
memcpy((void*)©->Map[0], (void*)&Map[0], MapCount * sizeof(Map[0]));
|
||||
copy->SpillSlotCount = SpillSlotCount;
|
||||
copy->MapCount = MapCount;
|
||||
copy->IsShared = IsShared;
|
||||
return UniquePtr {copy};
|
||||
}
|
||||
|
||||
|
||||
@@ -81,7 +81,7 @@ ssize_t LoadFileToBuffer(const fextl::string& Filepath, std::span<char> Buffer)
|
||||
#else
|
||||
template<typename T>
|
||||
static bool LoadFileImpl(T& Data, const fextl::string& Filepath, size_t FixedSize) {
|
||||
std::ifstream f(Filepath, std::ios::binary | std::ios::ate);
|
||||
std::ifstream f(Filepath.c_str(), std::ios::binary | std::ios::ate);
|
||||
if (f.fail()) {
|
||||
return false;
|
||||
}
|
||||
@@ -93,7 +93,7 @@ static bool LoadFileImpl(T& Data, const fextl::string& Filepath, size_t FixedSiz
|
||||
}
|
||||
|
||||
ssize_t LoadFileToBuffer(const fextl::string& Filepath, std::span<char> Buffer) {
|
||||
std::ifstream f(Filepath, std::ios::binary | std::ios::ate);
|
||||
std::ifstream f(Filepath.c_str(), std::ios::binary | std::ios::ate);
|
||||
return f.readsome(Buffer.data(), Buffer.size());
|
||||
}
|
||||
|
||||
|
||||
@@ -22,13 +22,13 @@ namespace FEXCore::Utils::SpinWaitLock {
|
||||
*/
|
||||
#ifdef _M_ARM_64
|
||||
|
||||
#define LOADEXCLUSIVE(LoadExclusiveOp, RegSize) \
|
||||
#define LOADEXCLUSIVE(LoadExclusiveOp, RegSize) \
|
||||
/* Prime the exclusive monitor with the passed in address. */ \
|
||||
#LoadExclusiveOp " %" #RegSize "[Result], [%[Futex]];"
|
||||
|
||||
#define SPINLOOP_BODY(LoadAtomicOp, RegSize) \
|
||||
#define SPINLOOP_BODY(LoadAtomicOp, RegSize) \
|
||||
/* WFE will wait for either the memory to change or spurious wake-up. */ \
|
||||
"wfe;" /* Load with acquire to get the result of memory. */ \
|
||||
"wfe;" /* Load with acquire to get the result of memory. */ \
|
||||
#LoadAtomicOp " %" #RegSize "[Result], [%[Futex]]; "
|
||||
|
||||
#define SPINLOOP_WFE_LDX_8BIT LOADEXCLUSIVE(ldaxrb, w)
|
||||
|
||||
@@ -47,6 +47,8 @@ uint64_t SetSignalMask(uint64_t Mask) {
|
||||
#ifndef _WIN32
|
||||
::syscall(SYS_rt_sigprocmask, SIG_SETMASK, &Mask, &Mask, 8);
|
||||
return Mask;
|
||||
#else
|
||||
return 0;
|
||||
#endif
|
||||
}
|
||||
|
||||
|
||||
@@ -0,0 +1,152 @@
|
||||
# What is x86-TSO and what is different compared to ARM's weak memory model?
|
||||
x86's memory model is a very strictly coherent memory model that effectively mandates that all memory accesses are "atomic". While atomicity is
|
||||
actually a bit more strict, we actually need to emulate it in ARM using atomic instructions. We are also required to emulate this strictness with
|
||||
unaligned accesses, which is due to x86 CPUs allowing unaligned atomics for "free" within a cacheline. Intel also takes this a step more and allowing
|
||||
full atomics with a feature called "split-locks", AMD gains this same feature in Zen 5.
|
||||
|
||||
# Emulating loads
|
||||
Due to x86 SIB addressing, this can happen on most instructions. FEX emulates these in a variety of ways depending on features.
|
||||
Most instructions are emulated with an atomic instruction but we also implement a feature called "half-barrier" atomics for unaligned atomics.
|
||||
|
||||
## Base ARMv8.0
|
||||
- Addressing limitations
|
||||
- Register only
|
||||
This is emulated using an atomic load instruction plus a nop.
|
||||
- On unaligned access the code gets backpatched to a non-atomic load plus a memory barrier
|
||||
|
||||
## FEAT_LRCPC
|
||||
- Addressing limitations
|
||||
- Register only
|
||||
This matches the base ARMv8.0 implementation but adds new instructions that match x86-TSO behaviour, making the emulation slightly quicker.
|
||||
- On unaligned access it still gets backpatched to non-atomic load plus a memory barrier.
|
||||
|
||||
## FEAT_LRCPC2
|
||||
- Addressing limitations
|
||||
- Register plus 9-bit signed immediate (-256, 255)
|
||||
Adds some new instructions that allow immediate encoding inside of the previous LRCPC instructions
|
||||
|
||||
## FEAT_LRCPC3
|
||||
Adds a handful of GPR instructions that aren't super interesting
|
||||
|
||||
FEX doesn't currently implement these since no hardware supports it.
|
||||
|
||||
- ldapr - Post-index load for stack
|
||||
- ldiapr - Post-index load pair for stack
|
||||
- stilp - pre-index store pair for stack
|
||||
- stlr - pre-index store for stack
|
||||
|
||||
# Emulating stores
|
||||
Again due to x86 SIB addressing, this can also happen on most instructions. There are less options for FEX with this extension, so in most cases this
|
||||
just turns in to an atomic store with half-barrier backpatching for unaligned accesses
|
||||
|
||||
## FEAT_LRCPC, FEAT_LRCPC2
|
||||
Adds nothing for emulating stores
|
||||
|
||||
## FEAT_LRCPC3
|
||||
|
||||
# Emulating atomic instructions
|
||||
x86 has atomic memory operations that can do a variety of operations. For unaligned atomic operations FEX will emulate the operation inside the signal
|
||||
handler if it happens to be unaligned.
|
||||
|
||||
## CASPair - cmpxchg
|
||||
|
||||
## Base ARMv8.0
|
||||
- Addressing limitations
|
||||
- Register only
|
||||
This is emulated with a ldaxp+stlxp pair of instructions.
|
||||
|
||||
## FEAT_LSE
|
||||
- Addressing limitations
|
||||
- Register only
|
||||
Adds a new caspal instruction that does the operation almost exactly like x86.
|
||||
|
||||
## CAS - cmpxchg8b/cmpxchg16b
|
||||
|
||||
## Base ARMv8.0
|
||||
- Addressing limitations
|
||||
- Register only
|
||||
Similar to CASPair but now only uses a ldaxr+stlxr pair
|
||||
|
||||
## FEAT_LSE
|
||||
- Addressing limitations
|
||||
- Register only
|
||||
Similar to CASPair adds a new casal instruction that operates basically like x86
|
||||
|
||||
# AtomicFetch<Op>
|
||||
## Op from the following list
|
||||
- Add
|
||||
- Sub
|
||||
- And
|
||||
- CLR
|
||||
- Or
|
||||
- Xor
|
||||
- Neg
|
||||
- Swap
|
||||
|
||||
## Base ARMv8.0
|
||||
- Addressing limitations
|
||||
- Register only
|
||||
All operations get emulated with an ldaxr+stlxr+<op> instruction
|
||||
|
||||
## FEAT_LSE
|
||||
- Addressing limitations
|
||||
- Register only
|
||||
Almost all operations now have a native atomic memory operation instruction. The only outlier is atomicNeg which doesn't have an LSE equivalent and
|
||||
uses the ARMv8.0 implementation.
|
||||
|
||||
# Vector loads
|
||||
Since almost all memory accesses on x86 are TSO, this includes vector operations.
|
||||
|
||||
## Base ARMv8.0
|
||||
- Addressing limitations
|
||||
- Register plus 9-bit signed immediate (-256, 255)
|
||||
- Register plus 12-bit unsigned scaled immediate (Scaled by access size)
|
||||
Emulated using half-barriers, which means a load+dmb
|
||||
|
||||
## FEAT_LRCPC3
|
||||
- LDAP1 added for element loads. Register only address encoding
|
||||
- LDAPUR added for vector register loads, supports 9-bit simm offset
|
||||
|
||||
# Vector stores
|
||||
Just like loads, these are emulated using half-barriers
|
||||
|
||||
## Base ARMv8.0
|
||||
- Addressing limitations
|
||||
- Register plus 9-bit signed immediate (-256, 255)
|
||||
- Register plus 12-bit unsigned scaled immediate (Scaled by access size)
|
||||
Emulated using half-barriers, which means a dmb+str
|
||||
|
||||
|
||||
## FEAT_LRCPC3
|
||||
- STL1 added for element stores. Register only address encoding
|
||||
- STLUR added for vector register stores, supports 9-bit simm offset
|
||||
|
||||
# Addressing limitations depending on operating mode
|
||||
## GPR loadstores
|
||||
### TSO Emulation disabled
|
||||
- Register only (ldr/str)
|
||||
- Register + Register + scale (ldr/str)
|
||||
- Register + 9-bit simm (ldur/stru)
|
||||
- Register + 12-bit unsigned scaled imm (ldr/str)
|
||||
|
||||
### TSO Emulation enabled
|
||||
- Register only (ldar/stlr)
|
||||
- Register only (ldapr/stlr) - FEAT_LRCPC
|
||||
- Register + 9-bit simm (ldapr/stlur) - FEAT_LRCPC2
|
||||
|
||||
## Vector loadstores
|
||||
### TSO Emulation disabled
|
||||
- Register only (ldr/str)
|
||||
- Register + Register + scale (ldr/str)
|
||||
- Register + 9-bit simm (ldur/stru)
|
||||
- Register + 12-bit unsigned scaled imm (ldr/str)
|
||||
|
||||
### TSO Emulation enabled
|
||||
- Same as TSO emulation disabled due to half-barrier implementation
|
||||
|
||||
### TSO Emulation enabled (FEAT_LRCPC3)
|
||||
- Register only (ldap1/stl1) - Element loadstore
|
||||
- Register + 9-bit simm (ldapur/stlur)
|
||||
|
||||
## Atomic memory operations
|
||||
Always TSO emulation enabled, always register only.
|
||||
Loaded 100 of 436 files, more files were not shown because too many files have changed in this diff.
Show more
Reference in new issue
Block a user